Acquisition, Query Forms, and Problem Extensions
An acquisition function is the rule that decides what to ask next (Section 12.1). In preferential Bayesian optimization (PBO) it chooses a pair, or a set, of options to show a person, and Section 19.3 and Section 19.4 introduced the main candidates. This chapter follows those rules through the research record from 2017 to September 2026: where each came from, what is proved about it, under which conditions it did well or badly, and what forms a query can take besides "A or B?".
The design of acquisition functions went through three phases: heuristic and information-theoretic rules from 2017 to 2021; a turn to decision theory in 2022 and 2023, when the expected utility of the best option (EUBO) and its multi-option form qEUBO were proved one-step Bayes optimal for noise-free answers; and, from 2024 to 2026, kernelized dueling algorithms driven by frequentist regret (Chapter 29) alongside criticism and repair of the EUBO default. Two findings frame everything below. The failure modes reported by different groups agree with one another. And the empirical comparisons were run in settings that cannot be compared with one another; at small budgets with realistic noise, random queries sometimes keep up.
28.1 2017 to 2021: heuristics and information #
Two rules from before 2017. The interactive Bayesian optimization of Brochu et al. (2007) chose as the first option the queried point with the largest posterior mean, and as the second the point with the largest expected improvement over it (Section 12.3); Astudillo et al. (2023) later noted that qEUBO with two options, if forced to include the current best point, reduces to exactly this rule. Bayesian active learning by disagreement (BALD), proposed for Gaussian process classifiers by Houlsby et al. (2011) (a preprint), chooses the query whose answer has the largest mutual information with the model's parameters (Section 6.3): the query on which plausible models disagree most. The dueling information gain of Benavoli et al. (2021c) extends BALD to preferences.
The three rules of González et al. González et al. (2017) proposed three acquisition functions on the dueling space (Section 19.2). Pure exploration (PE) picks the pair whose duel outcome has the largest variance. Copeland expected improvement (CEI) computes the one-step lookahead improvement in the soft-Copeland value of the Condorcet winner, the option that beats every other with probability above one half. Dueling Thompson sampling (DTS) picks the first point to maximize the soft-Copeland score computed from one continuous Thompson sample of the preference function (Section 12.5), and the second purely to explore, as the point whose duel against the first is most uncertain. In their experiments on one- and two-dimensional functions (Table 28.3), DTS was consistently the best strategy, and the dueling-bandit baseline Sparring needed about 4000 iterations to approach what DTS reached in 200 duels.
Duels among many, and easy questions. SelfSparring reduced multi-dueling, in which several options are compared in each round, to an ordinary bandit problem solved by Thompson sampling; its kernel version, KernelSelfSparring, adds a Gaussian process prior so that information is shared between options (Sui et al., 2017b). Bıyık et al. (2019) showed (their Theorem 1) that the global optimum of the commonly used volume removal objective, which scores a query by how much of the space of reward parameters an answer is expected to rule out, is a trivial query of identical options. Replacing it with information gain favors queries a person can answer with confidence, and learned faster in simulation and in a user study. This work belongs to preference-based reward learning with a parametric reward, not to Gaussian process PBO (Section 36.1).
A burst of new rules, 2020 and 2021. Projective PBO (Mikkola et al., 2020) came with five rules for its projective queries, among them projective expected improvement and preferential coordinate descent. Benavoli et al. (2021c) defined three rules relative to the current winner : a dueling upper confidence bound, the upper end of the 95% credible interval of ; dueling Thompson sampling; and EIIG, the logarithm of the probability of improvement plus times the dueling information gain, with or . Nguyen et al. (2021) proposed multinomial predictive entropy search (MPES), which they describe as the first information-theoretic acquisition function for Bayesian optimization with preference observations (Section 12.7); it optimizes all inputs of a query jointly, but it must enumerate the possible answers, so it suits only small query sets. Siivola et al. (2021) adapted batch expected improvement (qEI) and batch Thompson sampling to batch-winner feedback, in which the person names the best option of a set.
Separating two kinds of uncertainty. Fauvel and Chalk (2021) (a preprint) split the uncertainty about a duel's outcome into an epistemic part, which more data would remove, and an aleatoric part, the person's own randomness, which it would not. Their maximally uncertain challenge (MUC) takes as champion the point of largest posterior mean and as challenger the point that maximizes the epistemic variance of , in closed form and with a batch version. In the same year came the first kernelized dueling algorithm with a cumulative regret guarantee (Kirschner and Krause, 2021), whose feedback model Section 29.2 examines.
Sources cited in Section 28.1 12
- Brochu et al. (2007) Active Preference Learning with Discrete Choice Data
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Houlsby et al. (2011) Bayesian Active Learning for Classification and Preference Learning
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- González et al. (2017) Preferential Bayesian Optimization
- Sui et al. (2017b) Multi-dueling Bandits with Dependent Arms
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
- Kirschner and Krause (2021) Bias-Robust Bayesian Optimization via Dueling Bandits
28.2 2022 to 2023: EUBO and the decision-theoretic turn #
EUBO in preference exploration. EUBO first appeared in Bayesian optimization with preference exploration (BOPE), where experiments produce several outcomes and a decision maker's utility over those outcomes is learned from comparisons. Lin et al. (2022) used EUBO to choose the two outcome vectors to show the decision maker, and proved that it is the one-step Bayes optimal preference-exploration policy (Section 19.4.1). Two variants address the fact that some outcomes are not achievable: EUBO-ζ generates outcomes from one posterior sample of the outcome model, and EUBO-f̃ compares only outcomes that are likely achievable. On an unrestricted outcome set, EUBO tended to over-explore outcomes that could not be achieved. The paper also concluded that a Monte Carlo version of BALD, BALD-f̃, is a strong and fast baseline.
qEUBO. Astudillo et al. (2023) generalized EUBO to
where is the number of options shown to the person in one query, from which the person picks the best, and is the latent utility (written in Section 19.4). It is not a batch of separate queries. Section 19.4.1 summarized their four theorems; their conditions matter for how far they reach:
- Theorem 1. With noise-free answers, qEUBO is one-step Bayes optimal and equivalent to the knowledge gradient (Section 12.6).
- Theorem 2. Under logistic (multinomial logit) noise of scale (Equation (16.4); the paper writes ), the one-step value of qEUBO's maximizer is at least the noise-free one-step optimal value minus , where is the Lambert W function, the inverse of (Exercise 28.1 computes how fast the guarantee loosens with ).
- Theorem 3. On a finite domain, with and further technical conditions, the Bayesian simple regret of qEUBO, the expected gap between the best utility and that of the recommended option averaged over the prior (Section 13.1), is . Examples of sufficient conditions are a logistic likelihood with a prior under which, almost surely, for all ; or a nondegenerate Gaussian process prior with a likelihood equal to a constant whenever .
- Theorem 4. On some instances satisfying the same assumptions, the Bayesian simple regret of qEI stays above a constant for every : qEI is not asymptotically consistent.
Two readings follow (inference). EUBO's one-step optimality was first proved
in BOPE; qEUBO's contribution is the extension to noise and to , and
consistency on finite domains. And the conditions behind the rate, a
finite domain with utility gaps bounded away from zero, turn the problem into
one of identifying the best of finitely many options, so the rate cannot be
compared with the rates for continuous domains in Section 29.4. In
software, BoTorch 0.10.0 (February 2024) added qEUBO (Meta Platforms, Inc., 2026e),
and its preference tutorial runs the loop with PairwiseGP and the analytic
EUBO (Meta Platforms, Inc., 2026c).
The hallucination believer. In the same year, Takeno et al. (2023) proposed the hallucination believer (HB) of Section 19.3: take the current winner as the first point of the duel, and apply expected improvement or an upper confidence bound to a Gaussian process fitted to one sample of the latent comparison values from the truncated posterior (a hallucination). They had observed that standard acquisition functions applied to a preference Gaussian process keep choosing similar duels, because a duel carries little information and the variance of the preference model hardly decreases. Ignatenko et al. (2025) later proposed remaining system uncertainty, from a minimax view of data collection, as a performance measure that needs no ground truth.
Sources cited in Section 28.2 6
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
- Meta Platforms, Inc. (2026c) Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1)
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Ignatenko et al. (2025) On preference learning based on sequential Bayesian optimization with pairwise comparison
28.3 2024 to 2026: noise, exact knowledge gradient, amortization #
2024 and 2025. Ozaki et al. (2024) used a mutual-information active-learning acquisition function to choose preference queries in a multi-objective setting (see Section 28.6). Sinaga et al. (2026) (first posted in 2024; its 2026 version appeared in the proceedings track of ProbML 2026) proposed a risk-averse acquisition function that trades utility against how hard a comparison is to answer, and showed that their risk-adjusted EUBO stays one-step Bayes optimal up to an additive constant. POP-BO (Xu et al., 2024b) and MaxMinLCB (Pásztor et al., 2024) are optimistic algorithms with regret bounds (Section 29.3), and in 2025 PABBO (Zhang et al., 2025a) used a pretrained transformer policy that outputs query pairs directly.
2026: the default under examination. Most of the 2026 method work examined the default pipeline, in two preprints reported with the failure modes below (Wu and Gardner, 2026; Shao et al., 2026). Other work adapted classical ideas. Erarslan et al. (2025) (a preprint) adapted max-value entropy search to a production-cost constraint under which every comparison must involve a candidate that has already been produced. Haltia et al. (2026) (a preprint) used a cost-aware value of information to choose between a direct evaluation and a pairwise query; they report performance close to the convex hull of the two single-source trajectories, and the method falls back to standard Bayesian optimization when queries are expensive or noisy. Local PBO (Menn et al., 2026a) (a preprint) ported trust-region search (TuRPBO) and derivative-guided local search (GIPBO, PrefSQP) to pairwise feedback (Section 30.4). And PF-TS (Lazzaro et al., 2026) chooses a pair by drawing two independent posterior samples and maximizing each against a common anchor point.
Sources cited in Section 28.3 11
- Ozaki et al. (2024) Multi-Objective Bayesian Optimization with Active Preference Learning
- Sinaga et al. (2026) Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Wu and Gardner (2026) Knowledge Gradient for Preference Learning
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Erarslan et al. (2025) Consecutive Preferential Bayesian Optimization
- Haltia et al. (2026) Elicitation-Augmented Bayesian Optimization
- Menn et al. (2026a) Local Preferential Bayesian Optimization
- Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
28.4 Documented failure modes #
Section 19.6 previewed these failures. Here they are with the mechanism each paper gives and the conditions under which each was seen.
Adapted expected improvement stalls, again and again. Four groups reported the same phenomenon independently. González et al. found that CEI over-exploits and that the interactive Bayesian optimization of Brochu et al. performs poorly (González et al., 2017). Fauvel and Chalk found Brochu et al.'s expected improvement only slightly better than random, and attributed it to a frequent pathology in which the function samples the same duel members (Fauvel and Chalk, 2021). Takeno et al. found that expectation propagation with expected improvement often stalls through over-exploitation (Takeno et al., 2023). And Astudillo et al. found that qEI tended to stall late in a run, on the 7-dimensional Alpine1 function with initial data that included comparisons against a known good point, and proved the mechanism in their Theorem 4: when the value of the incumbent, the point of largest posterior mean, is already known fairly precisely, qEI is reluctant to include it in a query, so it learns only how the other options compare with one another (Astudillo et al., 2023). Rules of this family stop testing a well-known incumbent (inference; Section 19.3).
EUBO over-exploits, and its pipeline becomes ill-conditioned. On BOPE's unrestricted outcome set, EUBO over-explored unachievable outcomes (Lin et al., 2022). In single-objective PBO, EUBO's queries collapse toward the estimated maximum, as Section 19.6 reported and Section 19.5.1 reproduces in simulation. The source, Wu and Gardner (2026) (a preprint), adds the mechanism: under a Gaussian process prior and a probit likelihood, the one-step lookahead posterior is an extended skew-normal distribution, a skewed relative of the Gaussian whose mean has a closed form, and EUBO is only a lower bound on the exact knowledge gradient, approximately equal to it only when the probit noise goes to zero. Their test case is the two-dimensional Levy function, and their abstract also acknowledges a case showing that the knowledge gradient has limits in some situations. KappaSharp, also a preprint, reports that EUBO's queries create isolated comparison pairs and a rank-deficient Hessian (Shao et al., 2026) (Section 27.6). And POP-BO's authors report that, on instances sampled from a Gaussian process, qEUBO's reported solution was slightly better than theirs, but its cumulative regret was more than 2.5 times higher (Xu et al., 2024b).
Thompson sampling over-explores as the dimension grows. DTS was the best rule in one and two dimensions (González et al., 2017), but rules based on Thompson sampling performed only modestly over 34 functions, where KernelSelfSparring's weaker batch performance was attributed to its choosing the members of a batch independently (Fauvel and Chalk, 2021), and Thompson sampling over-explored on the 4- and 6-dimensional Hartmann functions (Takeno et al., 2023). Every study that tested a Thompson-sampling rule at four or more dimensions found it behind another rule (Figure 28.1). A plausible mechanism, which none of these papers isolates, is that the maximum of one posterior sample tends to fall where the posterior is most uncertain, and the share of the domain that is far from every observation grows quickly with dimension (inference; Exercise 28.2).
The hallucination believer depends on the noise. Takeno et al.'s results were obtained at noise variance , exactly the regime in which they showed the Laplace approximation fails worst. Their Appendix G.4 shows that the hallucination believer did relatively poorly when combined with the maximally uncertain challenge or with binary expected improvement, which they suspect is due to over-exploration (Takeno et al., 2023). Under logistic noise, the POP-BO authors report that the hallucination believer got stuck in local optima because it trusts random preference feedback too much, treating it as a hard constraint when it draws its Thompson sample (Xu et al., 2024b). Its advantage, in other words, depends on the noise level (inference).
Other costs. MPES must enumerate the possible answers: in qEUBO's 4- to 7-dimensional experiments it took 12.7 to 24.8 seconds per iteration against about 7 to 12 seconds for qEUBO, and it lost to qEUBO (Astudillo et al., 2023).
Sources cited in Section 28.4 8
- González et al. (2017) Preferential Bayesian Optimization
- Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Wu and Gardner (2026) Knowledge Gradient for Preference Learning
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
28.5 The acquisition functions compared #
Table 28.1 collects the main rules. Every entry under "Favorable evidence" and "Documented failures" carries the dimensions and noise under which it was observed, taken from the study that reported it (Table 28.3 gives each study's full setting). The frequentist regret bounds are detailed in Chapter 29.
| Rule (proposers, venue) | Theory | Favorable evidence | Documented failures |
|---|---|---|---|
| PE (González et al., 2017) | none | 1-D and 2-D, 33 grid points per dimension; noise not stated | worse as dimension grows |
| CEI (González et al., 2017) | none | run only on the 1-D Forrester function | over-exploits; too costly to compute |
| DTS (González et al., 2017) | no regret bound for one objective; a multi-objective version is asymptotically consistent (Astudillo et al., 2025) | best on 1-D and 2-D grids; noise not stated | over-explores from 4-D up (4-D and 6-D Hartmann, noise variance ) (Takeno et al., 2023); fourth of nine rules over 34 functions (probit, unit variance) (Fauvel and Chalk, 2021) |
| KernelSelfSparring (Sui et al., 2017b) | the independent-arm version is asymptotically no-regret; the kernel version only conjectured | no dedicated evidence | falls behind in batches because members are chosen independently (Fauvel and Chalk, 2021) |
| Adapted EI and qEI (Brochu et al., 2007; Siivola et al., 2021) | qEI not asymptotically consistent (qEUBO Theorem 4) | no clear difference from batch Thompson sampling in batch experiments (at most 4-D, utility noise sd 0.05) | stalls, over-exploits, picks identical duel members (four groups; 1-D to 7-D, noise from to logistic) |
| MUC (Fauvel and Chalk, 2021) | none | tied first by Borda rank over 34 functions (probit, unit variance) | relatively poor when combined with HB; stalled on Bukin and Ackley under expectation propagation |
| Dueling UCB, dueling Thompson sampling, EIIG (Benavoli et al., 2021c) | none | dueling UCB tied first over 34 functions (probit, unit variance) (Fauvel and Chalk, 2021) | EIIG seventh in the same comparison |
| MPES (Nguyen et al., 2021) | none | consistently best in 1-D to 3-D and on SUSHI; noise not stated | must enumerate answers; lost to qEUBO and slowest in 4-D to 7-D (logistic noise) (Astudillo et al., 2023) |
| BALD (Houlsby et al., 2011; Lin et al., 2022; Meta Platforms, Inc., 2026e) | none | competitive but slightly worse in BOPE (10% wrong choices) | none recorded specifically |
| EUBO, qEUBO (Lin et al., 2022; Astudillo et al., 2023) | one-step Bayes optimal without noise, equivalent to the knowledge gradient; additive-constant guarantee under logistic noise; Bayesian simple regret on finite domains with | best on all but one problem in 4-D to 7-D, moderate logistic noise, 150 queries | queries collapse toward the estimated maximum (2-D Levy, probit; preprint) (Wu and Gardner, 2026); rank-deficient Hessian (preprint) (Shao et al., 2026); higher cumulative regret (6-D Ackley and GP samples, logistic) (Xu et al., 2024b) |
| HB (Takeno et al., 2023) | none (listed as future work) | best overall over 12 functions up to 6-D at noise variance | stuck in local optima under logistic noise (Xu et al., 2024b); over-explores when combined with MUC |
| Exact knowledge gradient (Wu and Gardner, 2026) | EUBO is a lower bound on it | a 2-D Levy case study (probit) | its abstract acknowledges limits in some situations |
| POP-BO, MaxMinLCB, MR-LPF, PF-TS (Xu et al., 2024b; Pásztor et al., 2024; Kayal et al., 2025; Lazzaro et al., 2026) | see Section 29.4 | all low-dimensional experiments (1-D to 6-D, logistic) | PF-TS reports higher cumulative regret for MR-LPF at (1-D Ackley, 3-D catalyst) (Lazzaro et al., 2026) |
| PABBO (Zhang et al., 2025a) | none | first or second on most tasks (1-D, 2-D, 6-D, HPO-B, Candy, Sushi; noise-free) | fixed dimension; noise-free default evaluation; weaker on 6-D Hartmann |
The conditions in the table are easier to compare as a picture. Figure 28.1 places every comparison study by its noise model and the input dimensions it tested, and lets you choose a rule to see which studies report on it and with what result.
Some things to look for:
- Thompson sampling (the default view): both favorable findings sit at one to three dimensions, and even there it trailed MPES.
- The hallucination believer: favorable in the top row, near noise-free answers; unfavorable in the logistic-noise row.
- EUBO and qEUBO: favorable at 4 to 7 dimensions with logistic noise; mixed or unfavorable in the 2024 to 2026 studies that measured cumulative regret or looked for collapse.
- The empty regions: above 20 dimensions, only the local-PBO preprint compares rules.
Sources cited in Section 28.5 20
- González et al. (2017) Preferential Bayesian Optimization
- Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
- Sui et al. (2017b) Multi-dueling Bandits with Dependent Arms
- Brochu et al. (2007) Active Preference Learning with Discrete Choice Data
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Houlsby et al. (2011) Bayesian Active Learning for Classification and Preference Learning
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
- Wu and Gardner (2026) Knowledge Gradient for Preference Learning
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
- Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
- Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
28.6 Query forms #
A query need not be a pair. Chapter 20 teaches the forms and reports the evidence for the main ones. For larger sets (Section 20.1.2), qEUBO found clearly better than , while Siivola et al. found only marginal gains from larger batches, with different acquisition functions and noise (Astudillo et al., 2023; Siivola et al., 2021). For galleries and projections (Section 20.3), every projective variant beat every pairwise variant in 2 to 20 dimensions (Mikkola et al., 2020), and the Sequential Gallery's plane search beat line search in 5 to 20 dimensions, with 6 participants satisfied after 5.36 iterations on average (Koyama et al., 2020). With pairs alone, Figure 19.4 shows how quickly the share of the gap that forty comparisons close shrinks as parameters are added. Benavoli et al. (2023) let a person pick several mutually incomparable options from a set. This section adds the remaining forms and what the evidence on forms, taken together, says.
Coactive feedback, ordinal labels, and robot-specific forms. CoSpar (Tucker et al., 2020b) adds coactive feedback to the posterior sampling of SelfSparring: the user both compares trials and suggests improvements. LineCoSpar (Tucker et al., 2020a) restricts the computation to a random line (Section 20.3.3) and tuned 6 gait parameters with 6 able-bodied participants. ROIAL (Li et al., 2021) combines ordinal labels with preferences inside a region of interest meant to guarantee safety and comfort. Chapter 24 follows this line through a case study.
Preferences over hypothetical outcomes, and requests for improvement. Astudillo and Frazier (2020) let a decision maker compare attribute vectors, and BOPE alternates a stage in which the person compares outcome vectors that may be hypothetical, sampled from the outcome model, with an experimentation stage (Meta Platforms, Inc., 2026d). Ozaki et al. (2024) add an improvement request: the decision maker indicates which objective of a shown result they want improved, and the utility is a Chebyshev scalarization (a weighted worst-case combination of the objectives) with uncertain weights.
Weak preferences, validity labels, numbers, and language. The "About Equal"
answer of Bıyık et al. (2019), C-GLISp's better, worse, or similar, validity labels,
and crash reports are extensions of the observation model
(Section 27.2). Xu et al. (2020b) allowed both direct queries and duels
(COMP-GP-UCB). For finitely many options, Wang et al. (2025b) proved that an
efficient algorithm pays, for each option, only the smaller of the two regrets
it would incur from reward feedback or from dueling feedback. Social BO
(Adachi et al., 2025), a preprint, proved that under mild rationality axioms
noisy group feedback alone cannot reach a consensus free of social influence,
and mixed cheap public votes with expensive private ones. Natural language
enters as a source of labels: PEBOL (Austin et al., 2024a) uses
natural-language inference as its likelihood over independent items, and LILO
(Kobalczyk et al., 2026) has a language model translate free-text feedback into
pairwise labels for PairwiseGP with qEUBO, with a simulated decision maker and
no study with people (Section 35.2). Labels generated by a language model
are correlated and biased rather than independent probit noise (inference).
What the evidence on forms says. It points two ways. Forms that let each human action carry more information (projections, planes, ) beat pairs in the papers that introduced them, while the only direct comparison of batch winners and full rankings found little difference. Forms that ask for a continuous answer (sliders, projections, coactive suggestions) bring other kinds of noise, motor and perceptual precision and cognitive load, which every cited paper handles with a single generic noise term (inference). As of September 2026 we found no controlled study with people that compares pairs, batch winners, full rankings, and sliders under the same acquisition function and budget, and no acquisition function derived for forms other than pairs and multi-option sets: slider, plane, and projective queries use adapted expected improvement or random subspaces. Until such a study exists, the choice of form is best made on the human side, by what people can answer reliably and quickly (Section 20.3.4, Section 32.4), and a richer form should be tested against pairs within the same system (inference).
Sources cited in Section 28.6 17
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Benavoli et al. (2023) Learning Choice Functions with Gaussian Processes
- Tucker et al. (2020b) Preference-Based Learning for Exoskeleton Gait Optimization
- Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
- Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
- Astudillo and Frazier (2020) Multi-attribute Bayesian optimization with interactive preference learning
- Meta Platforms, Inc. (2026d) Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1)
- Ozaki et al. (2024) Multi-Objective Bayesian Optimization with Active Preference Learning
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
- Xu et al. (2020b) Zeroth Order Non-convex optimization with Dueling-Choice Bandits
- Wang et al. (2025b) Fusing Reward and Dueling Feedback in Stochastic Bandits
- Adachi et al. (2025) Bayesian Optimization for Building Social-Influence-Free Consensus
- Austin et al. (2024a) Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation
- Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback
28.7 Problem extensions #
PBO has been extended in the same directions as ordinary Bayesian optimization: several objectives, constraints, context, many dimensions, mixed inputs, several fidelities, and stopping. Table 28.2 summarizes each direction; the paragraphs after it give the points that matter.
| Extension | Representative work (in time order) | Type of evidence | Main gap |
|---|---|---|---|
| Several objectives and preferences over outcomes | Astudillo and Frazier 2020 (Astudillo and Frazier, 2020); BOPE 2022 (Lin et al., 2022); Ozaki et al. 2024 (Ozaki et al., 2024); PUB-MOBO 2025 (Ip et al., 2025); Astudillo et al. 2025 (Astudillo et al., 2025); Wang et al. 2025 (Wang et al., 2025a); Huber et al. 2025 (Huber et al., 2025); Active-MoSH 2026 (Chen et al., 2026) | simulated decision makers; engineering benchmarks | few studies with people |
| Constraints | StageOpt 2018 (Sui et al., 2018b); Benavoli et al. 2021 (Benavoli et al., 2021c); C-GLISp 2022 (Zhu et al., 2022); Kwon et al. 2022 (Kwon et al., 2022); Iwai et al. 2025 (Iwai et al., 2025) | controller calibration; 11 designers | no new constraint paper after 2025 |
| Safety | StageOpt 2018 (Sui et al., 2018b); ROIAL 2021 (Li et al., 2021); Cosner et al. 2022 (Cosner et al., 2022); CrashPBO 2026 (Menn et al., 2026b) | spinal cord stimulation; a quadruped robot; three robot platforms | none named |
| Context | Khan et al. 2025 (Khan et al., 2025); Wang et al. 2025 (Wang et al., 2025d); Coutinho et al. 2025, 2026 (Coutinho et al., 2025; Coutinho et al., 2026) | simulated buildings; a report of negative transfer | multi-task Gaussian processes untested with a preference likelihood |
| Transfer across users | Granley et al. 2023 (Granley et al., 2023); Meta-PO 2025 (Li et al., 2025a); PABBO 2025 (Zhang et al., 2025a) | 36 people for Meta-PO | few hierarchical utility priors |
| High dimension | sequential line search 2017 (Koyama et al., 2017); projective PBO 2020 (Mikkola et al., 2020); LineCoSpar 2020 (Tucker et al., 2020a); Sequential Gallery 2020 (Koyama et al., 2020); qEUBO 2023 (Astudillo et al., 2023); local methods 2026 (Menn et al., 2026a); GimmBO 2026 (Liu et al., 2026b) | simulated comparisons up to 102 dimensions | dimension-scaled and sparse priors untested |
| Mixed and categorical inputs | piecewise affine surrogates 2025 (Zhu and Bemporad, 2025) | benchmarks | no Gaussian process preference method with categorical kernels |
| Several fidelities | Theiner et al. 2026 (Theiner et al., 2026); elicitation-augmented BO 2026 (Haltia et al., 2026) | preprint or conference paper | none before mid-2025 |
| Stopping | Bıyık et al. 2019, Theorem 3 (Bıyık et al., 2019); the Sequential Gallery's satisfaction button (Koyama et al., 2020); Ignatenko et al. 2025 (Ignatenko et al., 2025) | parametric models | no stopping rule for pairwise Gaussian processes |
Several objectives. Astudillo et al. (2025) let every objective be observed only through preferences and proposed dueling scalarized Thompson sampling (DSTS): sample from the posterior, apply a random Chebyshev scalarization, then run dueling Thompson sampling. They proved it asymptotically consistent, which they call the first convergence guarantee for dueling Thompson sampling in PBO, while noting that even for single-objective PBO the regret bound of DTS remains unknown; DSTS did best on four synthetic functions and on simulated exoskeleton and autonomous-driving tasks. The "preferences" of Abdolshah et al. (2019) (NeurIPS 2019) are an importance order over objectives, not pairwise feedback on designs; the other works in the table continue the line of Section 14.5.
Constraints and safety. StageOpt (Sui et al., 2018b) optimizes a utility under unknown safety constraints and separates expanding the safe region from maximizing utility. Its guarantees of constraint satisfaction and convergence are stated for numerical observations; a variant in an appendix learns the utility from preference feedback while the safety functions still receive numerical measurements, comes without a convergence theorem of its own, and was used clinically for spinal cord stimulation. Constrained PBO (Iwai et al., 2025) proposed EUBOC, which weights EUBO by the probability, modeled with a Gaussian process, that the constraints are satisfied, in the manner of constrained expected improvement (Gardner et al., 2014), and evaluated it in a banner-ad study with 11 professional designers, with predicted click-through rate as the constraint. Constraints have been handled in two ways: by learning feasibility from labels a person provides (C-GLISp, Benavoli et al.'s validity labels, crash feedback), or by measuring the constraint separately (the click-through rate in constrained PBO, StageOpt's safety signal). Which route fits depends on whether the constraint can be observed without the person (inference).
Context and transfer. Context enters through offline utilities learned from expert knowledge (Khan et al., 2025) or through contextual variables such as outdoor temperature in controller tuning (Wang et al., 2025d), a preprint. Transfer has been implemented through amortization (PABBO), stored models of earlier users (Meta-PO), and encoders trained on a population (Granley et al.), not through a multi-task Gaussian process across users (inference); Section 32.5 reports what population priors buy.
High dimension, mixed inputs, fidelities, and stopping. The high-dimensional methods of 2017 to 2020 all restrict each query to a low-dimensional subspace through the current best point, which turns a -dimensional acquisition optimization into a one- or two-dimensional one and lets each human action carry more than one bit (inference); Section 30.3.1 lists how far each reached, and Section 30.4 examines the local methods of 2026 and the lengthscale bound that confounds their comparison with qEUBO. The piecewise affine surrogate of Zhu and Bemporad (2025) uses mixed-integer linear programming to handle known linear constraints and mixed numerical and categorical variables; it is the only preference method for mixed inputs we found, and the Sequential Gallery lists not handling discrete parameters, such as layouts, fonts, or filter types, among its limitations. The only multi-fidelity preference methods are those of Theiner et al. (2026) and elicitation-augmented Bayesian optimization (Haltia et al., 2026), both from 2026. For stopping, the only explicit optimal rule assumes a parametric reward model (Bıyık et al., 2019), and the Sequential Gallery stops when the user presses a "satisfied" button; Section 30.7 reports what scalar Bayesian optimization has learned about stopping and what a preferential rule would need.
Sources cited in Section 28.7 37
- Astudillo and Frazier (2020) Multi-attribute Bayesian optimization with interactive preference learning
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Ozaki et al. (2024) Multi-Objective Bayesian Optimization with Active Preference Learning
- Ip et al. (2025) User Preference Meets Pareto-Optimality in Multi-Objective Bayesian Optimization
- Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
- Wang et al. (2025a) Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble
- Huber et al. (2025) Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization
- Chen et al. (2026) Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
- Sui et al. (2018b) Stagewise Safe Bayesian Optimization with Gaussian Processes
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
- Kwon et al. (2022) Physically Consistent Preferential Bayesian Optimization for Food Arrangement
- Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
- Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
- Cosner et al. (2022) Safety-Aware Preference-Based Learning for Safety-Critical Control
- Menn et al. (2026b) Preferential Bayesian Optimization with Crash Feedback
- Khan et al. (2025) Efficient Contextual Preferential Bayesian Optimization with Historical Examples
- Wang et al. (2025d) Personalized Building Climate Control with Contextual Preferential Bayesian Optimization
- Coutinho et al. (2025) Accelerated controller tuning using human feedback and Multi-Task Preferential Bayesian Optimization
- Coutinho et al. (2026) Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization
- Granley et al. (2023) Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses
- Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Menn et al. (2026a) Local Preferential Bayesian Optimization
- Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
- Zhu and Bemporad (2025) Global and Preference-Based Optimization with Mixed Variables Using Piecewise Affine Surrogates
- Theiner et al. (2026) Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models
- Haltia et al. (2026) Elicitation-Augmented Bayesian Optimization
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
- Ignatenko et al. (2025) On preference learning based on sequential Bayesian optimization with pairwise comparison
- Abdolshah et al. (2019) Multi-objective Bayesian optimisation with preferences over objectives
- Gardner et al. (2014) Bayesian Optimization with Inequality Constraints
28.8 Competing 'first' claims #
PBO sits between dueling bandits, preference-based reinforcement learning, control (the GLISp line), and human-computer interaction, and claims of being "first" in one community often overlook earlier work in another. Each claim below holds only within its exact setting (inference).
- Constrained PBO (Iwai et al., 2025) claims to be the first to introduce inequality constraints, after C-GLISp (Zhu et al., 2022) handled unknown constraints in 2022 and a variant of StageOpt (Sui et al., 2018b) learned a utility from preference feedback under numerically measured safety constraints in 2018; it does not cite the validity labels of Benavoli et al. 2021 (Benavoli et al., 2021c). The claim holds for inequality constraints on a Gaussian process surrogate.
- LineSpar (Cheng et al., 2020), a workshop paper, claims to be the first high-dimensional preference-based Bayesian optimization, after sequential line search (2017) and alongside projective PBO (ICML 2020).
- DSTS (Astudillo et al., 2025) claims the first multi-objective PBO framework, which overlaps with a preprint on choice functions (Benavoli et al., 2021b); the claim holds for latent objectives observed only through preferences.
- POP-BO (Xu et al., 2024b) says existing methods lack cumulative regret or global convergence guarantees, with the qualifier "continuous input space", and does not cite Kirschner and Krause (2021).
- MPES (Nguyen et al., 2021) calls itself the first information-theoretic acquisition function with preference observations (first arXiv version December 2020), while the EIIG rule (Benavoli et al., 2021c) (August 2020) also uses the dueling information gain; both descend from BALD (Houlsby et al., 2011) (inference).
- EUBO's one-step Bayes optimality was first shown in BOPE (Lin et al., 2022); the qEUBO paper's claim holds for its extension to logistic noise and .
Sources cited in Section 28.8 12
- Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
- Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
- Sui et al. (2018b) Stagewise Safe Bayesian Optimization with Gaussian Processes
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- Cheng et al. (2020) Preference-Based Bayesian Optimization in High Dimensions with Human Feedback
- Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
- Benavoli et al. (2021b) Choice functions based multi-objective Bayesian optimisation
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Kirschner and Krause (2021) Bias-Robust Bayesian Optimization via Dueling Bandits
- Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
- Houlsby et al. (2011) Bayesian Active Learning for Classification and Preference Learning
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
28.9 The state of empirical comparison #
Table 28.3 lists the main comparison studies and their settings. In it, qTS is batch Thompson sampling, qNEI is noisy batch expected improvement, HPO-B is a benchmark built from hyperparameter optimization tasks, and a Sobol sequence is a quasi-random space-filling design.
| Study | Compared | Dimensions | Budget and repetitions | Noise model | Main conclusion |
|---|---|---|---|---|---|
| González et al. 2017 (González et al., 2017) | PE, CEI, DTS, random, interactive BO, Sparring | 1, 2 | 5 + 200 duels; 20 repetitions | not stated | DTS best |
| Mikkola et al. 2020 (Mikkola et al., 2020) | 5 projective rules against pairwise random and DTS variants | 2, 6, 10, 20 | 100 queries; 25 initializations | small Gaussian noise | every projective variant beat every pairwise variant |
| Siivola et al. 2021 (Siivola et al., 2021) | batch EI, batch Thompson sampling, random; three inference methods | 6 functions; real data with | batch sizes 2 to 6; 10 repetitions | utility noise sd 0.05 | differences between acquisition functions larger than between feedback types; only slightly better than baseline on real data |
| Nguyen et al. 2021 (Nguyen et al., 2021) | MPES, EI, DTS | 1 to 3; CIFAR-10 embedding; SUSHI | not stated | not stated | MPES consistently best |
| Fauvel and Chalk 2021 (Fauvel and Chalk, 2021) | 9 rules | 34 functions | 80 iterations; 40 repetitions | probit, unit variance after normalization | MUC and dueling UCB tied first; Brochu EI eighth; random ninth |
| Lin et al. 2022 (Lin et al., 2022) | EUBO-ζ, EUBO-f̃, BALD-f̃, random, others | multi-outcome problems | 75 comparisons in 3 stages; 30 repetitions | 10% wrong choices | EUBO variants best |
| Astudillo et al. 2023 (Astudillo et al., 2023) | qEUBO, MPES, qTS, qEI, qNEI, random | 4 to 7 | initial + 150 queries; 50 or 100 repetitions | logistic, calibrated to 10%, 20%, 30% errors on the top 1% of point pairs | with , qEUBO best on all problems except Car cab |
| Takeno et al. 2023 (Takeno et al., 2023) | HB, Laplace and EP with EI, MUC, skew-GP sampling rules | up to 6 (12 functions) | initial; 10 repetitions | noise variance | HB best overall; qEUBO not included |
| Xu et al. 2024, POP-BO (Xu et al., 2024b) | DTS, HB, qEUBO | GP samples; 6-D Ackley | not stated | logistic | qEUBO's reported solution slightly better, cumulative regret more than 2.5 times higher |
| Zhang et al. 2025, PABBO (Zhang et al., 2025a) | qEUBO, qEI, qNEI, qTS, MPES, random | 1, 2, 6; HPO-B; Candy; Sushi | 30 repetitions | noise-free | PABBO first or second; random often beat some GP baselines |
| Lazzaro et al. 2026, PF-TS (Lazzaro et al., 2026) | MR-LPF, POP-BO, MaxMinLCB | 1-D Ackley; 3-D catalyst | ; 30 repetitions | logistic | PF-TS cumulative regret significantly lower than MR-LPF and POP-BO |
| Menn et al. 2026, local (preprint) (Menn et al., 2026a) | local methods, qEUBO, HB with EI, GLISp, Sobol | up to 102 | about comparisons, after random evaluations for policy search | Gaussian, 10% of the value range | local methods better on steep optima |
The noise calibration of the qEUBO experiments comes from their Appendix C.2 and code (Astudillo, 2023b).
The comparisons do not add up. The studies differ in metric (Section 31.4.3 lists five definitions of regret), in noise model, from a nearly noise-free probit to 10% flipped answers and moderate Gumbel noise (Section 31.4.2), and in posterior inference: Laplace, expectation propagation, variational inference, or skew Gaussian process sampling. Rankings flip with these choices, as POP-BO's comparison with qEUBO shows (inference). Fauvel and Chalk's Borda analysis over 34 functions is the broadest single comparison, but it predates qEUBO and the hallucination believer, uses expectation propagation throughout, and runs only 80 iterations. Takeno et al.'s comparison comes next, but its duels are almost noise-free and it does not include qEUBO. No comparison has breadth, realistic noise, and the acquisition functions of 2023 and later at the same time.
Random queries sometimes keep up. The PABBO authors write that the random strategy often beat some of the Gaussian process baselines (Zhang et al., 2025a). On real data with a low signal-to-noise ratio, Siivola et al. found that all methods only barely beat the baseline, and wrote that preference observations are inherently less informative than direct ones, that larger batches do not alleviate this, and that for very noisy data, random search with large batches may be a good choice (Siivola et al., 2021). In a single replication, BoTorch's BOPE tutorial prints candidate utilities of for EUBO-ζ, for random preference exploration, for the true utility, and for random experimentation, and notes that EUBO-ζ's win is not guaranteed in a single replication (Meta Platforms, Inc., 2026d). The three pieces of evidence point the same way: at small budgets and realistic noise, the gap between an elaborate acquisition function and a random design may be small, and a single run says little (inference).
Evidence from people. The human studies in this chapter are all small, and none was designed to compare acquisition functions: projective PBO's materials-science user, the Sequential Gallery's 6 participants, LineCoSpar's 6, constrained PBO's 11 designers, and Meta-PO's 36. As of September 2026, we found no study that randomizes people to different acquisition functions, such as qEUBO, DTS, and MUC, under the same interface and budget. Section 47.4 describes the experiment that would settle it. Until then, a practitioner should treat the ranking of acquisition functions as unknown for people and keep a share of random queries in every session as a control, which costs little given the evidence above (inference). Nor is there a benchmark for PBO with agreed metrics, noise models, and budgets: each paper reuses Forrester, six-hump camel, Hartmann, Ackley, Sushi, and Candy as it sees fit (Section 31.5).
Sources cited in Section 28.9 14
- González et al. (2017) Preferential Bayesian Optimization
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
- Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
- Menn et al. (2026a) Local Preferential Bayesian Optimization
- Astudillo (2023b) qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py)
- Meta Platforms, Inc. (2026d) Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1)
28.10 Settled, contested, missing #
Settled. qEI is not asymptotically consistent on the instances the qEUBO paper constructs (Astudillo et al., 2023). Adapted expected improvement stalled in the experiments of four independent groups. EUBO's one-step Bayes optimality holds without noise, has an additive-constant guarantee under logistic noise, and was first shown in BOPE (Lin et al., 2022). Query forms in which each action carries more information, projections and planes, beat pairs in the simulations of the papers that proposed them (Mikkola et al., 2020; Koyama et al., 2020).
Contested. The value of more options per query: qEUBO found clearly better than , Siivola et al. found only marginal gains, and their acquisition functions and noise differ. The advantage of the hallucination believer: it is best near noise-free answers, and stuck in local optima under logistic noise according to POP-BO. EUBO's collapse and ill-conditioning: two 2026 preprints, not yet peer reviewed or tested on human data. The ranking of Thompson-sampling rules, which flips with dimension and setting.
Missing. A broad comparison of acquisition functions on a shared benchmark with matched noise and budget. An experiment that randomizes people to acquisition functions. Acquisition functions derived for slider, plane, and projective queries. A stopping rule with guarantees for pairwise Gaussian process models. Regret bounds for DTS and HB, and continuous-domain guarantees for qEUBO under noise (Chapter 29).
For a choice that must be made now. The evidence supports including random
queries and simple baselines among the comparisons, and reporting dimension,
noise, and metric. qEUBO with PairwiseGP is the best-maintained default, but
that recommendation comes from the software ecosystem and from within the PBO
framework, not from a direct comparison with simpler methods (inference).
Sources cited in Section 28.10 4
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
28.11 Exercises #
Theorem 2 of the qEUBO paper bounds the loss from logistic noise of scale by , where is the inverse of . Compute for and by solving , and say what the bound implies about showing more options per query when answers are noisy.
Solution
For , solve : gives , so . For , solve : gives , so . The guaranteed loss relative to the noise-free one-step optimum roughly doubles from to , in units of the noise scale , while the value of a noise-free answer from four options is at least that from two. The bound alone does not say whether more options help; it says the guarantee loosens slowly, which is consistent with qEUBO's experiments finding better than at moderate noise.
A rough way to see why Thompson sampling over-explores as the dimension grows. Suppose the posterior is confident within distance of each of observed points in the unit cube , and uncertain elsewhere. Estimate the fraction of the cube that is confident for and , using the volume of a -dimensional ball, , and ignoring overlaps and edges. Where will the maximum of one posterior sample tend to fall?
Solution
For the "ball" is an interval of length , so 20 points could cover the whole interval (the estimate exceeds 1; overlaps make the true coverage at most 1). For , the ball volume is , so 20 balls cover about of the cube. Almost all of the six-dimensional cube is uncertain, so a posterior sample has many chances to be high somewhere nobody has looked, and its maximum tends to land there. That is exploration by construction, and in six dimensions it rarely returns to refine the best region. This picture is our illustration of the mechanism, not a result of the cited papers, which report the over-exploration without isolating its cause.
Using Table 28.3, find a study that compares qEUBO with the hallucination believer. Under which noise model, at which dimensions, and with which metric? What would a comparison need in order to settle whether either rule is better at human noise levels?
Solution
Only the POP-BO paper (Xu et al., 2024b) includes both (with DTS), under logistic noise, on Gaussian process samples and the 6-dimensional Ackley function, with an unstated budget. It reports that the hallucination believer got stuck in local optima, and that qEUBO's reported solution was slightly better than POP-BO's but its cumulative regret more than 2.5 times higher. Takeno et al., where the hallucination believer did best, did not include qEUBO and used noise variance . A settling comparison would fix a noise model calibrated to human comparisons, run both rules with the same surrogate, inference, budget, and dimensions, report both simple and cumulative regret, include random queries, and ideally be repeated with people.
Sources cited in Section 28.11 1
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
Further reading #
- Astudillo et al. (2023) is the reference for qEUBO; read its four theorems with their conditions, and its Appendix C.2 on how the noise was calibrated.
- Fauvel and Chalk (2021) is the broadest comparison of rules (34 functions, nine rules) and the clearest treatment of epistemic against aleatoric uncertainty in acquisition.
- Takeno et al. (2023) and Xu et al. (2024b) disagree about the hallucination believer; together they show how much the noise level decides.
- Wu and Gardner (2026) derives the exact knowledge gradient under a probit likelihood and is the best place to see why EUBO's equivalence with it breaks under noise.
- Mikkola et al. (2020) and Koyama et al. (2020) are the two reference papers on queries richer than a pair.
References
- (2019). Multi-objective Bayesian optimisation with preferences over objectives. Advances in Neural Information Processing Systems. Cited in §28.7
- (2025). Bayesian Optimization for Building Social-Influence-Free Consensus. arXiv. preprint Cited in §28.6
- (2023b). qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py). GitHub. software Cited in §28.9
- (2020). Multi-attribute Bayesian optimization with interactive preference learning. International Conference on Artificial Intelligence and Statistics. Cited in §28.6 §28.7
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §28.1 §28.2 §28.4 §28.5 §28.6 §28.7 §28.9 §28.10
- (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Cited in §28.5 §28.7 §28.8
- (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Cited in §28.6
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Cited in §28.1 §28.6 §28.7
- (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Cited in §28.1 §28.5
- (2026). Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds. Conference on Uncertainty in Artificial Intelligence. Cited in §28.7
- (2020). Preference-Based Bayesian Optimization in High Dimensions with Human Feedback. SCMLS 2020 Workshop. workshop paper Cited in §28.8
- (2022). Safety-Aware Preference-Based Learning for Safety-Critical Control. Learning for Dynamics and Control Conference. Cited in §28.7
- (2025). Accelerated controller tuning using human feedback and Multi-Task Preferential Bayesian Optimization. 2025 American Control Conference (ACC). Cited in §28.7
- (2026). Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization. Control Engineering Practice. Cited in §28.7
- (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Cited in §28.3
- (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. preprint Cited in §28.1 §28.4 §28.5 §28.9
- (2014). Bayesian Optimization with Inequality Constraints. Proceedings of the 31st International Conference on Machine Learning (ICML 2014). Cited in §28.7
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §28.1 §28.4 §28.5 §28.9
- (2023). Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. NeurIPS 2023. Cited in §28.7
- (2026). Elicitation-Augmented Bayesian Optimization. arXiv. preprint Cited in §28.3 §28.7
- (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Cited in §28.1 §28.5 §28.8
- (2025). Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization. Journal of Multi-Criteria Decision Analysis. Cited in §28.7
- (2025). On preference learning based on sequential Bayesian optimization with pairwise comparison. Artificial Intelligence. Cited in §28.2 §28.7
- (2025). User Preference Meets Pareto-Optimality in Multi-Objective Bayesian Optimization. Proceedings of the AAAI Conference on Artificial Intelligence. Cited in §28.7
- (2025). Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. Cited in §28.7 §28.8
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §28.5
- (2025). Efficient Contextual Preferential Bayesian Optimization with Historical Examples. Proceedings of the Genetic and Evolutionary Computation Conference Companion. Cited in §28.7
- (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. Cited in §28.1 §28.8
- (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Cited in §28.6
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Cited in §28.7
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Cited in §28.6 §28.7 §28.10
- (2022). Physically Consistent Preferential Bayesian Optimization for Food Arrangement. IEEE Robotics and Automation Letters. Cited in §28.7
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §28.3 §28.5 §28.9
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Cited in §28.6 §28.7
- (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Cited in §28.7
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §28.2 §28.4 §28.5 §28.7 §28.8 §28.9 §28.10
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §28.7
- (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Cited in §28.3 §28.7 §28.9
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Cited in §28.7
- (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. software Cited in §28.2
- (2026d). Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1). botorch.org. software Cited in §28.6 §28.9
- (2026e). BoTorch CHANGELOG. GitHub. software Cited in §28.2 §28.5
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §28.1 §28.6 §28.7 §28.9 §28.10
- (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Cited in §28.1 §28.5 §28.8 §28.9
- (2024). Multi-Objective Bayesian Optimization with Active Preference Learning. Proceedings of the AAAI Conference on Artificial Intelligence. Cited in §28.3 §28.6 §28.7
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Cited in §28.3 §28.5
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §28.3 §28.4 §28.5
- (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Cited in §28.1 §28.5 §28.6 §28.9
- (2026). Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization. Symposium on Probabilistic Machine Learning (ProbML 2026), Proceedings Track. Cited in §28.3
- (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. Cited in §28.1 §28.5
- (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. Cited in §28.7 §28.8
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §28.2 §28.4 §28.5 §28.9
- (2026). Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models. European Control Conference. Cited in §28.7
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Cited in §28.6 §28.7
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Cited in §28.6
- (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. Cited in §28.7
- (2025b). Fusing Reward and Dueling Feedback in Stochastic Bandits. International Conference on Machine Learning. Cited in §28.6
- (2025d). Personalized Building Climate Control with Contextual Preferential Bayesian Optimization. arXiv. preprint Cited in §28.7
- (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Cited in §28.3 §28.4 §28.5
- (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. Cited in §28.6
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §28.3 §28.4 §28.5 §28.8 §28.9 §28.11
- (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §28.3 §28.5 §28.7 §28.9
- (2025). Global and Preference-Based Optimization with Mixed Variables Using Piecewise Affine Surrogates. Journal of Optimization Theory and Applications. Cited in §28.7
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Cited in §28.7 §28.8