Bayesian Optimization
Part VI: The Research Frontier
中文

Acquisition, Query Forms, and Problem Extensions

An acquisition function is the rule that decides what to ask next (Section 12.1). In preferential Bayesian optimization (PBO) it chooses a pair, or a set, of options to show a person, and Section 19.3 and Section 19.4 introduced the main candidates. This chapter follows those rules through the research record from 2017 to September 2026: where each came from, what is proved about it, under which conditions it did well or badly, and what forms a query can take besides "A or B?".

The design of acquisition functions went through three phases: heuristic and information-theoretic rules from 2017 to 2021; a turn to decision theory in 2022 and 2023, when the expected utility of the best option (EUBO) and its multi-option form qEUBO were proved one-step Bayes optimal for noise-free answers; and, from 2024 to 2026, kernelized dueling algorithms driven by frequentist regret (Chapter 29) alongside criticism and repair of the EUBO default. Two findings frame everything below. The failure modes reported by different groups agree with one another. And the empirical comparisons were run in settings that cannot be compared with one another; at small budgets with realistic noise, random queries sometimes keep up.

28.1 2017 to 2021: heuristics and information #

Two rules from before 2017. The interactive Bayesian optimization of Brochu et al. (2007) chose as the first option the queried point with the largest posterior mean, and as the second the point with the largest expected improvement over it (Section 12.3); Astudillo et al. (2023) later noted that qEUBO with two options, if forced to include the current best point, reduces to exactly this rule. Bayesian active learning by disagreement (BALD), proposed for Gaussian process classifiers by Houlsby et al. (2011) (a preprint), chooses the query whose answer has the largest mutual information with the model's parameters (Section 6.3): the query on which plausible models disagree most. The dueling information gain of Benavoli et al. (2021c) extends BALD to preferences.

The three rules of González et al. González et al. (2017) proposed three acquisition functions on the dueling space (Section 19.2). Pure exploration (PE) picks the pair whose duel outcome has the largest variance. Copeland expected improvement (CEI) computes the one-step lookahead improvement in the soft-Copeland value of the Condorcet winner, the option that beats every other with probability above one half. Dueling Thompson sampling (DTS) picks the first point to maximize the soft-Copeland score computed from one continuous Thompson sample of the preference function (Section 12.5), and the second purely to explore, as the point whose duel against the first is most uncertain. In their experiments on one- and two-dimensional functions (Table 28.3), DTS was consistently the best strategy, and the dueling-bandit baseline Sparring needed about 4000 iterations to approach what DTS reached in 200 duels.

Duels among many, and easy questions. SelfSparring reduced multi-dueling, in which several options are compared in each round, to an ordinary bandit problem solved by Thompson sampling; its kernel version, KernelSelfSparring, adds a Gaussian process prior so that information is shared between options (Sui et al., 2017b). Bıyık et al. (2019) showed (their Theorem 1) that the global optimum of the commonly used volume removal objective, which scores a query by how much of the space of reward parameters an answer is expected to rule out, is a trivial query of KK identical options. Replacing it with information gain favors queries a person can answer with confidence, and learned faster in simulation and in a user study. This work belongs to preference-based reward learning with a parametric reward, not to Gaussian process PBO (Section 36.1).

A burst of new rules, 2020 and 2021. Projective PBO (Mikkola et al., 2020) came with five rules for its projective queries, among them projective expected improvement and preferential coordinate descent. Benavoli et al. (2021c) defined three rules relative to the current winner xr\vx_r: a dueling upper confidence bound, the upper end of the 95% credible interval of f(x)−f(xr)f(\vx) - f(\vx_r); dueling Thompson sampling; and EIIG, the logarithm of the probability of improvement plus kk times the dueling information gain, with k=0.1k = 0.1 or 0.50.5. Nguyen et al. (2021) proposed multinomial predictive entropy search (MPES), which they describe as the first information-theoretic acquisition function for Bayesian optimization with preference observations (Section 12.7); it optimizes all inputs of a query jointly, but it must enumerate the possible answers, so it suits only small query sets. Siivola et al. (2021) adapted batch expected improvement (qEI) and batch Thompson sampling to batch-winner feedback, in which the person names the best option of a set.

Separating two kinds of uncertainty. Fauvel and Chalk (2021) (a preprint) split the uncertainty about a duel's outcome into an epistemic part, which more data would remove, and an aleatoric part, the person's own randomness, which it would not. Their maximally uncertain challenge (MUC) takes as champion the point of largest posterior mean and as challenger the point that maximizes the epistemic variance of Φ(f(x1)−f(x))\Phi(f(\vx_1) - f(\vx)), in closed form and with a batch version. In the same year came the first kernelized dueling algorithm with a cumulative regret guarantee (Kirschner and Krause, 2021), whose feedback model Section 29.2 examines.

Sources cited in Section 28.1 12
  1. Brochu et al. (2007) Active Preference Learning with Discrete Choice Data
  2. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  3. Houlsby et al. (2011) Bayesian Active Learning for Classification and Preference Learning
  4. Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
  5. González et al. (2017) Preferential Bayesian Optimization
  6. Sui et al. (2017b) Multi-dueling Bandits with Dependent Arms
  7. Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
  8. Mikkola et al. (2020) Projective Preferential Bayesian Optimization
  9. Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
  10. Siivola et al. (2021) Preferential Batch Bayesian Optimization
  11. Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
  12. Kirschner and Krause (2021) Bias-Robust Bayesian Optimization via Dueling Bandits

28.2 2022 to 2023: EUBO and the decision-theoretic turn #

EUBO in preference exploration. EUBO first appeared in Bayesian optimization with preference exploration (BOPE), where experiments produce several outcomes and a decision maker's utility over those outcomes is learned from comparisons. Lin et al. (2022) used EUBO to choose the two outcome vectors to show the decision maker, and proved that it is the one-step Bayes optimal preference-exploration policy (Section 19.4.1). Two variants address the fact that some outcomes are not achievable: EUBO-ζ generates outcomes from one posterior sample of the outcome model, and EUBO-f̃ compares only outcomes that are likely achievable. On an unrestricted outcome set, EUBO tended to over-explore outcomes that could not be achieved. The paper also concluded that a Monte Carlo version of BALD, BALD-f̃, is a strong and fast baseline.

qEUBO. Astudillo et al. (2023) generalized EUBO to

qEUBO⁡n(x1,…,xq)=En ⁣[max⁡{f(x1),…,f(xq)}],\operatorname{qEUBO}_n(\vx_1, \dots, \vx_q) = \E_n\!\left[\max\{f(\vx_1), \dots, f(\vx_q)\}\right],
(28.1)

where qq is the number of options shown to the person in one query, from which the person picks the best, and ff is the latent utility (written gg in Section 19.4). It is not a batch of separate queries. Section 19.4.1 summarized their four theorems; their conditions matter for how far they reach:

  1. Theorem 1. With noise-free answers, qEUBO is one-step Bayes optimal and equivalent to the knowledge gradient (Section 12.6).
  2. Theorem 2. Under logistic (multinomial logit) noise of scale τ\tau (Equation (16.4); the paper writes λ\lambda), the one-step value of qEUBO's maximizer is at least the noise-free one-step optimal value minus τ W ⁣((q−1)/e)\tau\, W\!\left((q - 1)/e\right), where WW is the Lambert W function, the inverse of w↦weww \mapsto w e^w (Exercise 28.1 computes how fast the guarantee loosens with qq).
  3. Theorem 3. On a finite domain, with q=2q = 2 and further technical conditions, the Bayesian simple regret of qEUBO, the expected gap between the best utility and that of the recommended option averaged over the prior (Section 13.1), is o(1/n)o(1/n). Examples of sufficient conditions are a logistic likelihood with a prior under which, almost surely, δ≤∣f(x)−f(y)∣≤Δ\delta \le \lvert f(\vx) - f(\vy)\rvert \le \Delta for all x≠y\vx \neq \vy; or a nondegenerate Gaussian process prior with a likelihood equal to a constant a>1/2a > 1/2 whenever f(x1)≠f(x2)f(\vx_1) \neq f(\vx_2).
  4. Theorem 4. On some instances satisfying the same assumptions, the Bayesian simple regret of qEI stays above a constant R>0R > 0 for every nn: qEI is not asymptotically consistent.

Two readings follow (inference). EUBO's one-step optimality was first proved in BOPE; qEUBO's contribution is the extension to noise and to q>2q > 2, and consistency on finite domains. And the conditions behind the o(1/n)o(1/n) rate, a finite domain with utility gaps bounded away from zero, turn the problem into one of identifying the best of finitely many options, so the rate cannot be compared with the rates for continuous domains in Section 29.4. In software, BoTorch 0.10.0 (February 2024) added qEUBO (Meta Platforms, Inc., 2026e), and its preference tutorial runs the loop with PairwiseGP and the analytic EUBO (Meta Platforms, Inc., 2026c).

The hallucination believer. In the same year, Takeno et al. (2023) proposed the hallucination believer (HB) of Section 19.3: take the current winner as the first point of the duel, and apply expected improvement or an upper confidence bound to a Gaussian process fitted to one sample of the latent comparison values from the truncated posterior (a hallucination). They had observed that standard acquisition functions applied to a preference Gaussian process keep choosing similar duels, because a duel carries little information and the variance of the preference model hardly decreases. Ignatenko et al. (2025) later proposed remaining system uncertainty, from a minimax view of data collection, as a performance measure that needs no ground truth.

Sources cited in Section 28.2 6
  1. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  2. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  3. Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
  4. Meta Platforms, Inc. (2026c) Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1)
  5. Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
  6. Ignatenko et al. (2025) On preference learning based on sequential Bayesian optimization with pairwise comparison

28.3 2024 to 2026: noise, exact knowledge gradient, amortization #

2024 and 2025. Ozaki et al. (2024) used a mutual-information active-learning acquisition function to choose preference queries in a multi-objective setting (see Section 28.6). Sinaga et al. (2026) (first posted in 2024; its 2026 version appeared in the proceedings track of ProbML 2026) proposed a risk-averse acquisition function that trades utility against how hard a comparison is to answer, and showed that their risk-adjusted EUBO stays one-step Bayes optimal up to an additive constant. POP-BO (Xu et al., 2024b) and MaxMinLCB (Pásztor et al., 2024) are optimistic algorithms with regret bounds (Section 29.3), and in 2025 PABBO (Zhang et al., 2025a) used a pretrained transformer policy that outputs query pairs directly.

2026: the default under examination. Most of the 2026 method work examined the default pipeline, in two preprints reported with the failure modes below (Wu and Gardner, 2026; Shao et al., 2026). Other work adapted classical ideas. Erarslan et al. (2025) (a preprint) adapted max-value entropy search to a production-cost constraint under which every comparison must involve a candidate that has already been produced. Haltia et al. (2026) (a preprint) used a cost-aware value of information to choose between a direct evaluation and a pairwise query; they report performance close to the convex hull of the two single-source trajectories, and the method falls back to standard Bayesian optimization when queries are expensive or noisy. Local PBO (Menn et al., 2026a) (a preprint) ported trust-region search (TuRPBO) and derivative-guided local search (GIPBO, PrefSQP) to pairwise feedback (Section 30.4). And PF-TS (Lazzaro et al., 2026) chooses a pair by drawing two independent posterior samples and maximizing each against a common anchor point.

Sources cited in Section 28.3 11
  1. Ozaki et al. (2024) Multi-Objective Bayesian Optimization with Active Preference Learning
  2. Sinaga et al. (2026) Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization
  3. Xu et al. (2024b) Principled Preferential Bayesian Optimization
  4. Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
  5. Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
  6. Wu and Gardner (2026) Knowledge Gradient for Preference Learning
  7. Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
  8. Erarslan et al. (2025) Consecutive Preferential Bayesian Optimization
  9. Haltia et al. (2026) Elicitation-Augmented Bayesian Optimization
  10. Menn et al. (2026a) Local Preferential Bayesian Optimization
  11. Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback

28.4 Documented failure modes #

Section 19.6 previewed these failures. Here they are with the mechanism each paper gives and the conditions under which each was seen.

Adapted expected improvement stalls, again and again. Four groups reported the same phenomenon independently. González et al. found that CEI over-exploits and that the interactive Bayesian optimization of Brochu et al. performs poorly (González et al., 2017). Fauvel and Chalk found Brochu et al.'s expected improvement only slightly better than random, and attributed it to a frequent pathology in which the function samples the same duel members (Fauvel and Chalk, 2021). Takeno et al. found that expectation propagation with expected improvement often stalls through over-exploitation (Takeno et al., 2023). And Astudillo et al. found that qEI tended to stall late in a run, on the 7-dimensional Alpine1 function with initial data that included comparisons against a known good point, and proved the mechanism in their Theorem 4: when the value of the incumbent, the point of largest posterior mean, is already known fairly precisely, qEI is reluctant to include it in a query, so it learns only how the other options compare with one another (Astudillo et al., 2023). Rules of this family stop testing a well-known incumbent (inference; Section 19.3).

EUBO over-exploits, and its pipeline becomes ill-conditioned. On BOPE's unrestricted outcome set, EUBO over-explored unachievable outcomes (Lin et al., 2022). In single-objective PBO, EUBO's queries collapse toward the estimated maximum, as Section 19.6 reported and Section 19.5.1 reproduces in simulation. The source, Wu and Gardner (2026) (a preprint), adds the mechanism: under a Gaussian process prior and a probit likelihood, the one-step lookahead posterior is an extended skew-normal distribution, a skewed relative of the Gaussian whose mean has a closed form, and EUBO is only a lower bound on the exact knowledge gradient, approximately equal to it only when the probit noise goes to zero. Their test case is the two-dimensional Levy function, and their abstract also acknowledges a case showing that the knowledge gradient has limits in some situations. KappaSharp, also a preprint, reports that EUBO's queries create isolated comparison pairs and a rank-deficient Hessian (Shao et al., 2026) (Section 27.6). And POP-BO's authors report that, on instances sampled from a Gaussian process, qEUBO's reported solution was slightly better than theirs, but its cumulative regret was more than 2.5 times higher (Xu et al., 2024b).

Thompson sampling over-explores as the dimension grows. DTS was the best rule in one and two dimensions (González et al., 2017), but rules based on Thompson sampling performed only modestly over 34 functions, where KernelSelfSparring's weaker batch performance was attributed to its choosing the members of a batch independently (Fauvel and Chalk, 2021), and Thompson sampling over-explored on the 4- and 6-dimensional Hartmann functions (Takeno et al., 2023). Every study that tested a Thompson-sampling rule at four or more dimensions found it behind another rule (Figure 28.1). A plausible mechanism, which none of these papers isolates, is that the maximum of one posterior sample tends to fall where the posterior is most uncertain, and the share of the domain that is far from every observation grows quickly with dimension (inference; Exercise 28.2).

The hallucination believer depends on the noise. Takeno et al.'s results were obtained at noise variance 10−410^{-4}, exactly the regime in which they showed the Laplace approximation fails worst. Their Appendix G.4 shows that the hallucination believer did relatively poorly when combined with the maximally uncertain challenge or with binary expected improvement, which they suspect is due to over-exploration (Takeno et al., 2023). Under logistic noise, the POP-BO authors report that the hallucination believer got stuck in local optima because it trusts random preference feedback too much, treating it as a hard constraint when it draws its Thompson sample (Xu et al., 2024b). Its advantage, in other words, depends on the noise level (inference).

Other costs. MPES must enumerate the possible answers: in qEUBO's 4- to 7-dimensional experiments it took 12.7 to 24.8 seconds per iteration against about 7 to 12 seconds for qEUBO, and it lost to qEUBO (Astudillo et al., 2023).

Sources cited in Section 28.4 8
  1. González et al. (2017) Preferential Bayesian Optimization
  2. Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
  3. Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
  4. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  5. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  6. Wu and Gardner (2026) Knowledge Gradient for Preference Learning
  7. Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
  8. Xu et al. (2024b) Principled Preferential Bayesian Optimization

28.5 The acquisition functions compared #

Table 28.1 collects the main rules. Every entry under "Favorable evidence" and "Documented failures" carries the dimensions and noise under which it was observed, taken from the study that reported it (Table 28.3 gives each study's full setting). The frequentist regret bounds are detailed in Chapter 29.

Table 28.1 The main acquisition rules for preferential Bayesian optimization, with the conditions under which each finding was observed.
Rule (proposers, venue) Theory Favorable evidence Documented failures
PE (González et al., 2017) none 1-D and 2-D, 33 grid points per dimension; noise not stated worse as dimension grows
CEI (González et al., 2017) none run only on the 1-D Forrester function over-exploits; too costly to compute
DTS (González et al., 2017) no regret bound for one objective; a multi-objective version is asymptotically consistent (Astudillo et al., 2025) best on 1-D and 2-D grids; noise not stated over-explores from 4-D up (4-D and 6-D Hartmann, noise variance 10−410^{-4}) (Takeno et al., 2023); fourth of nine rules over 34 functions (probit, unit variance) (Fauvel and Chalk, 2021)
KernelSelfSparring (Sui et al., 2017b) the independent-arm version is asymptotically no-regret; the kernel version only conjectured no dedicated evidence falls behind in batches because members are chosen independently (Fauvel and Chalk, 2021)
Adapted EI and qEI (Brochu et al., 2007; Siivola et al., 2021) qEI not asymptotically consistent (qEUBO Theorem 4) no clear difference from batch Thompson sampling in batch experiments (at most 4-D, utility noise sd 0.05) stalls, over-exploits, picks identical duel members (four groups; 1-D to 7-D, noise from 10−410^{-4} to logistic)
MUC (Fauvel and Chalk, 2021) none tied first by Borda rank over 34 functions (probit, unit variance) relatively poor when combined with HB; stalled on Bukin and Ackley under expectation propagation
Dueling UCB, dueling Thompson sampling, EIIG (Benavoli et al., 2021c) none dueling UCB tied first over 34 functions (probit, unit variance) (Fauvel and Chalk, 2021) EIIG seventh in the same comparison
MPES (Nguyen et al., 2021) none consistently best in 1-D to 3-D and on SUSHI; noise not stated must enumerate answers; lost to qEUBO and slowest in 4-D to 7-D (logistic noise) (Astudillo et al., 2023)
BALD (Houlsby et al., 2011; Lin et al., 2022; Meta Platforms, Inc., 2026e) none competitive but slightly worse in BOPE (10% wrong choices) none recorded specifically
EUBO, qEUBO (Lin et al., 2022; Astudillo et al., 2023) one-step Bayes optimal without noise, equivalent to the knowledge gradient; additive-constant guarantee under logistic noise; Bayesian simple regret o(1/n)o(1/n) on finite domains with q=2q = 2 best on all but one problem in 4-D to 7-D, moderate logistic noise, 150 queries queries collapse toward the estimated maximum (2-D Levy, probit; preprint) (Wu and Gardner, 2026); rank-deficient Hessian (preprint) (Shao et al., 2026); higher cumulative regret (6-D Ackley and GP samples, logistic) (Xu et al., 2024b)
HB (Takeno et al., 2023) none (listed as future work) best overall over 12 functions up to 6-D at noise variance 10−410^{-4} stuck in local optima under logistic noise (Xu et al., 2024b); over-explores when combined with MUC
Exact knowledge gradient (Wu and Gardner, 2026) EUBO is a lower bound on it a 2-D Levy case study (probit) its abstract acknowledges limits in some situations
POP-BO, MaxMinLCB, MR-LPF, PF-TS (Xu et al., 2024b; Pásztor et al., 2024; Kayal et al., 2025; Lazzaro et al., 2026) see Section 29.4 all low-dimensional experiments (1-D to 6-D, logistic) PF-TS reports higher cumulative regret for MR-LPF at T=300T = 300 (1-D Ackley, 3-D catalyst) (Lazzaro et al., 2026)
PABBO (Zhang et al., 2025a) none first or second on most tasks (1-D, 2-D, 6-D, HPO-B, Candy, Sushi; noise-free) fixed dimension; noise-free default evaluation; weaker on 6-D Hartmann

The conditions in the table are easier to compare as a picture. Figure 28.1 places every comparison study by its noise model and the input dimensions it tested, and lets you choose a rule to see which studies report on it and with what result.

not givenNoise-free or nearlyGaussian noise on the utilityProbit or logistic answersA share of wrong answersNoise model not statedZhang 2025Takeno 20235Mikkola 20207Siivola 20214Menn 2026Fauvel 20218Wu 2026Astudillo 20236Xu 2024Lazzaro 20262Lin 2022González 20171Nguyen 20213Koyama 2020125102050100input dimension (log scale)favorablemixedunfavorablelower dimensions also tested1González et al. 2017 · 1, 2-D · not statedDTS was consistently the best strategy.2Lazzaro et al. 2026 (PF-TS) · 1, 3-D · logisticPF-TS had significantly lower cumulative regret than MR-LPF and POP-BO.3Nguyen et al. 2021 (MPES) · 1 to 3-D, plus CIFAR-10 embedding, SUSHI · not statedDTS was behind MPES.4Siivola et al. 2021 · up to 4-D · utility noise sd 0.05Batch Thompson sampling and batch EI showed no clear difference.5Takeno et al. 2023 · up to 6-D · noise variance 10⁻⁴Thompson sampling over-explored on the 4-D and 6-D Hartmann functions.6Astudillo et al. 2023 (qEUBO) · 4 to 7-D · logistic, calibrated to 10%, 20%, 30% errors on the top1% of pairsqTS lost to qEUBO, which was best on every problem except Car cab (q = 2).7Mikkola et al. 2020 · 2, 6, 10, 20-D · small Gaussian noiseThe pairwise DTS variant lost to every projective variant, a contrast of query forms more thanof rules.8Fauvel and Chalk 2021 (preprint) · 34 functions, dimension not given · probit, unit variance afternormalizationDTS ranked fourth of nine rules; Thompson-based rules performed only modestly.
not givenNoise-free or nearlyGaussian noise on the utilityProbit or logistic answersA share of wrong answersNoise model not statedZhang 2025Takeno 20235Mikkola 20207Siivola 20214Menn 2026Fauvel 20218Wu 2026Astudillo 20236Xu 2024Lazzaro 20262Lin 2022González 20171Nguyen 20213Koyama 2020125102050100input dimension (log scale)favorablemixedunfavorablelower dimensions also tested1González et al. 2017 · 1, 2-D · not statedDTS was consistently the best strategy.2Lazzaro et al. 2026 (PF-TS) · 1, 3-D · logisticPF-TS had significantly lower cumulativeregret than MR-LPF and POP-BO.3Nguyen et al. 2021 (MPES) · 1 to 3-D, plusCIFAR-10 embedding, SUSHI · not statedDTS was behind MPES.4Siivola et al. 2021 · up to 4-D · utility noisesd 0.05Batch Thompson sampling and batch EI showedno clear difference.5Takeno et al. 2023 · up to 6-D · noise variance10⁻⁴Thompson sampling over-explored on the 4-Dand 6-D Hartmann functions.6Astudillo et al. 2023 (qEUBO) · 4 to 7-D ·logistic, calibrated to 10%, 20%, 30% errors onthe top 1% of pairsqTS lost to qEUBO, which was best on everyproblem except Car cab (q = 2).7Mikkola et al. 2020 · 2, 6, 10, 20-D · smallGaussian noiseThe pairwise DTS variant lost to everyprojective variant, a contrast of query formsmore than of rules.8Fauvel and Chalk 2021 (preprint) · 34functions, dimension not given · probit, unitvariance after normalizationDTS ranked fourth of nine rules;Thompson-based rules performed only modestly.
Figure 28.1 Where each finding about an acquisition rule was observed. Each row is a study, grouped by its noise model and drawn across the input dimensions it tested (log scale); a dashed segment means the study also tested lower dimensions without listing them, and a square in the right-hand column marks tasks whose dimension the study does not give, or real-data tasks. Choosing a rule colors the studies that report on it by whether the finding was favorable, mixed, or unfavorable for that rule, and lists the findings with their conditions. Findings and conditions are from the studies cited in Table 28.1 and Table 28.3; the verdict colors are this chapter's summary.

Some things to look for:

  • Thompson sampling (the default view): both favorable findings sit at one to three dimensions, and even there it trailed MPES.
  • The hallucination believer: favorable in the top row, near noise-free answers; unfavorable in the logistic-noise row.
  • EUBO and qEUBO: favorable at 4 to 7 dimensions with logistic noise; mixed or unfavorable in the 2024 to 2026 studies that measured cumulative regret or looked for collapse.
  • The empty regions: above 20 dimensions, only the local-PBO preprint compares rules.
Sources cited in Section 28.5 20
  1. González et al. (2017) Preferential Bayesian Optimization
  2. Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
  3. Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
  4. Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
  5. Sui et al. (2017b) Multi-dueling Bandits with Dependent Arms
  6. Brochu et al. (2007) Active Preference Learning with Discrete Choice Data
  7. Siivola et al. (2021) Preferential Batch Bayesian Optimization
  8. Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
  9. Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
  10. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  11. Houlsby et al. (2011) Bayesian Active Learning for Classification and Preference Learning
  12. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  13. Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
  14. Wu and Gardner (2026) Knowledge Gradient for Preference Learning
  15. Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
  16. Xu et al. (2024b) Principled Preferential Bayesian Optimization
  17. Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
  18. Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
  19. Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
  20. Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization

28.6 Query forms #

A query need not be a pair. Chapter 20 teaches the forms and reports the evidence for the main ones. For larger sets (Section 20.1.2), qEUBO found q=4q = 4 clearly better than q=2q = 2, while Siivola et al. found only marginal gains from larger batches, with different acquisition functions and noise (Astudillo et al., 2023; Siivola et al., 2021). For galleries and projections (Section 20.3), every projective variant beat every pairwise variant in 2 to 20 dimensions (Mikkola et al., 2020), and the Sequential Gallery's plane search beat line search in 5 to 20 dimensions, with 6 participants satisfied after 5.36 iterations on average (Koyama et al., 2020). With pairs alone, Figure 19.4 shows how quickly the share of the gap that forty comparisons close shrinks as parameters are added. Benavoli et al. (2023) let a person pick several mutually incomparable options from a set. This section adds the remaining forms and what the evidence on forms, taken together, says.

Coactive feedback, ordinal labels, and robot-specific forms. CoSpar (Tucker et al., 2020b) adds coactive feedback to the posterior sampling of SelfSparring: the user both compares trials and suggests improvements. LineCoSpar (Tucker et al., 2020a) restricts the computation to a random line (Section 20.3.3) and tuned 6 gait parameters with 6 able-bodied participants. ROIAL (Li et al., 2021) combines ordinal labels with preferences inside a region of interest meant to guarantee safety and comfort. Chapter 24 follows this line through a case study.

Preferences over hypothetical outcomes, and requests for improvement. Astudillo and Frazier (2020) let a decision maker compare attribute vectors, and BOPE alternates a stage in which the person compares outcome vectors that may be hypothetical, sampled from the outcome model, with an experimentation stage (Meta Platforms, Inc., 2026d). Ozaki et al. (2024) add an improvement request: the decision maker indicates which objective of a shown result they want improved, and the utility is a Chebyshev scalarization (a weighted worst-case combination of the objectives) with uncertain weights.

Weak preferences, validity labels, numbers, and language. The "About Equal" answer of Bıyık et al. (2019), C-GLISp's better, worse, or similar, validity labels, and crash reports are extensions of the observation model (Section 27.2). Xu et al. (2020b) allowed both direct queries and duels (COMP-GP-UCB). For finitely many options, Wang et al. (2025b) proved that an efficient algorithm pays, for each option, only the smaller of the two regrets it would incur from reward feedback or from dueling feedback. Social BO (Adachi et al., 2025), a preprint, proved that under mild rationality axioms noisy group feedback alone cannot reach a consensus free of social influence, and mixed cheap public votes with expensive private ones. Natural language enters as a source of labels: PEBOL (Austin et al., 2024a) uses natural-language inference as its likelihood over independent items, and LILO (Kobalczyk et al., 2026) has a language model translate free-text feedback into pairwise labels for PairwiseGP with qEUBO, with a simulated decision maker and no study with people (Section 35.2). Labels generated by a language model are correlated and biased rather than independent probit noise (inference).

What the evidence on forms says. It points two ways. Forms that let each human action carry more information (projections, planes, q=4q = 4) beat pairs in the papers that introduced them, while the only direct comparison of batch winners and full rankings found little difference. Forms that ask for a continuous answer (sliders, projections, coactive suggestions) bring other kinds of noise, motor and perceptual precision and cognitive load, which every cited paper handles with a single generic noise term (inference). As of September 2026 we found no controlled study with people that compares pairs, batch winners, full rankings, and sliders under the same acquisition function and budget, and no acquisition function derived for forms other than pairs and multi-option sets: slider, plane, and projective queries use adapted expected improvement or random subspaces. Until such a study exists, the choice of form is best made on the human side, by what people can answer reliably and quickly (Section 20.3.4, Section 32.4), and a richer form should be tested against pairs within the same system (inference).

Sources cited in Section 28.6 17
  1. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  2. Siivola et al. (2021) Preferential Batch Bayesian Optimization
  3. Mikkola et al. (2020) Projective Preferential Bayesian Optimization
  4. Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
  5. Benavoli et al. (2023) Learning Choice Functions with Gaussian Processes
  6. Tucker et al. (2020b) Preference-Based Learning for Exoskeleton Gait Optimization
  7. Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
  8. Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
  9. Astudillo and Frazier (2020) Multi-attribute Bayesian optimization with interactive preference learning
  10. Meta Platforms, Inc. (2026d) Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1)
  11. Ozaki et al. (2024) Multi-Objective Bayesian Optimization with Active Preference Learning
  12. Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
  13. Xu et al. (2020b) Zeroth Order Non-convex optimization with Dueling-Choice Bandits
  14. Wang et al. (2025b) Fusing Reward and Dueling Feedback in Stochastic Bandits
  15. Adachi et al. (2025) Bayesian Optimization for Building Social-Influence-Free Consensus
  16. Austin et al. (2024a) Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation
  17. Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback

28.7 Problem extensions #

PBO has been extended in the same directions as ordinary Bayesian optimization: several objectives, constraints, context, many dimensions, mixed inputs, several fidelities, and stopping. Table 28.2 summarizes each direction; the paragraphs after it give the points that matter.

Table 28.2 Extensions of preferential Bayesian optimization, the evidence behind each, and the main gap.
Extension Representative work (in time order) Type of evidence Main gap
Several objectives and preferences over outcomes Astudillo and Frazier 2020 (Astudillo and Frazier, 2020); BOPE 2022 (Lin et al., 2022); Ozaki et al. 2024 (Ozaki et al., 2024); PUB-MOBO 2025 (Ip et al., 2025); Astudillo et al. 2025 (Astudillo et al., 2025); Wang et al. 2025 (Wang et al., 2025a); Huber et al. 2025 (Huber et al., 2025); Active-MoSH 2026 (Chen et al., 2026) simulated decision makers; engineering benchmarks few studies with people
Constraints StageOpt 2018 (Sui et al., 2018b); Benavoli et al. 2021 (Benavoli et al., 2021c); C-GLISp 2022 (Zhu et al., 2022); Kwon et al. 2022 (Kwon et al., 2022); Iwai et al. 2025 (Iwai et al., 2025) controller calibration; 11 designers no new constraint paper after 2025
Safety StageOpt 2018 (Sui et al., 2018b); ROIAL 2021 (Li et al., 2021); Cosner et al. 2022 (Cosner et al., 2022); CrashPBO 2026 (Menn et al., 2026b) spinal cord stimulation; a quadruped robot; three robot platforms none named
Context Khan et al. 2025 (Khan et al., 2025); Wang et al. 2025 (Wang et al., 2025d); Coutinho et al. 2025, 2026 (Coutinho et al., 2025; Coutinho et al., 2026) simulated buildings; a report of negative transfer multi-task Gaussian processes untested with a preference likelihood
Transfer across users Granley et al. 2023 (Granley et al., 2023); Meta-PO 2025 (Li et al., 2025a); PABBO 2025 (Zhang et al., 2025a) 36 people for Meta-PO few hierarchical utility priors
High dimension sequential line search 2017 (Koyama et al., 2017); projective PBO 2020 (Mikkola et al., 2020); LineCoSpar 2020 (Tucker et al., 2020a); Sequential Gallery 2020 (Koyama et al., 2020); qEUBO 2023 (Astudillo et al., 2023); local methods 2026 (Menn et al., 2026a); GimmBO 2026 (Liu et al., 2026b) simulated comparisons up to 102 dimensions dimension-scaled and sparse priors untested
Mixed and categorical inputs piecewise affine surrogates 2025 (Zhu and Bemporad, 2025) benchmarks no Gaussian process preference method with categorical kernels
Several fidelities Theiner et al. 2026 (Theiner et al., 2026); elicitation-augmented BO 2026 (Haltia et al., 2026) preprint or conference paper none before mid-2025
Stopping Bıyık et al. 2019, Theorem 3 (Bıyık et al., 2019); the Sequential Gallery's satisfaction button (Koyama et al., 2020); Ignatenko et al. 2025 (Ignatenko et al., 2025) parametric models no stopping rule for pairwise Gaussian processes

Several objectives. Astudillo et al. (2025) let every objective be observed only through preferences and proposed dueling scalarized Thompson sampling (DSTS): sample from the posterior, apply a random Chebyshev scalarization, then run dueling Thompson sampling. They proved it asymptotically consistent, which they call the first convergence guarantee for dueling Thompson sampling in PBO, while noting that even for single-objective PBO the regret bound of DTS remains unknown; DSTS did best on four synthetic functions and on simulated exoskeleton and autonomous-driving tasks. The "preferences" of Abdolshah et al. (2019) (NeurIPS 2019) are an importance order over objectives, not pairwise feedback on designs; the other works in the table continue the line of Section 14.5.

Constraints and safety. StageOpt (Sui et al., 2018b) optimizes a utility under unknown safety constraints and separates expanding the safe region from maximizing utility. Its guarantees of constraint satisfaction and convergence are stated for numerical observations; a variant in an appendix learns the utility from preference feedback while the safety functions still receive numerical measurements, comes without a convergence theorem of its own, and was used clinically for spinal cord stimulation. Constrained PBO (Iwai et al., 2025) proposed EUBOC, which weights EUBO by the probability, modeled with a Gaussian process, that the constraints are satisfied, in the manner of constrained expected improvement (Gardner et al., 2014), and evaluated it in a banner-ad study with 11 professional designers, with predicted click-through rate as the constraint. Constraints have been handled in two ways: by learning feasibility from labels a person provides (C-GLISp, Benavoli et al.'s validity labels, crash feedback), or by measuring the constraint separately (the click-through rate in constrained PBO, StageOpt's safety signal). Which route fits depends on whether the constraint can be observed without the person (inference).

Context and transfer. Context enters through offline utilities learned from expert knowledge (Khan et al., 2025) or through contextual variables such as outdoor temperature in controller tuning (Wang et al., 2025d), a preprint. Transfer has been implemented through amortization (PABBO), stored models of earlier users (Meta-PO), and encoders trained on a population (Granley et al.), not through a multi-task Gaussian process across users (inference); Section 32.5 reports what population priors buy.

High dimension, mixed inputs, fidelities, and stopping. The high-dimensional methods of 2017 to 2020 all restrict each query to a low-dimensional subspace through the current best point, which turns a dd-dimensional acquisition optimization into a one- or two-dimensional one and lets each human action carry more than one bit (inference); Section 30.3.1 lists how far each reached, and Section 30.4 examines the local methods of 2026 and the lengthscale bound that confounds their comparison with qEUBO. The piecewise affine surrogate of Zhu and Bemporad (2025) uses mixed-integer linear programming to handle known linear constraints and mixed numerical and categorical variables; it is the only preference method for mixed inputs we found, and the Sequential Gallery lists not handling discrete parameters, such as layouts, fonts, or filter types, among its limitations. The only multi-fidelity preference methods are those of Theiner et al. (2026) and elicitation-augmented Bayesian optimization (Haltia et al., 2026), both from 2026. For stopping, the only explicit optimal rule assumes a parametric reward model (Bıyık et al., 2019), and the Sequential Gallery stops when the user presses a "satisfied" button; Section 30.7 reports what scalar Bayesian optimization has learned about stopping and what a preferential rule would need.

Sources cited in Section 28.7 37
  1. Astudillo and Frazier (2020) Multi-attribute Bayesian optimization with interactive preference learning
  2. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  3. Ozaki et al. (2024) Multi-Objective Bayesian Optimization with Active Preference Learning
  4. Ip et al. (2025) User Preference Meets Pareto-Optimality in Multi-Objective Bayesian Optimization
  5. Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
  6. Wang et al. (2025a) Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble
  7. Huber et al. (2025) Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization
  8. Chen et al. (2026) Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
  9. Sui et al. (2018b) Stagewise Safe Bayesian Optimization with Gaussian Processes
  10. Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
  11. Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
  12. Kwon et al. (2022) Physically Consistent Preferential Bayesian Optimization for Food Arrangement
  13. Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
  14. Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
  15. Cosner et al. (2022) Safety-Aware Preference-Based Learning for Safety-Critical Control
  16. Menn et al. (2026b) Preferential Bayesian Optimization with Crash Feedback
  17. Khan et al. (2025) Efficient Contextual Preferential Bayesian Optimization with Historical Examples
  18. Wang et al. (2025d) Personalized Building Climate Control with Contextual Preferential Bayesian Optimization
  19. Coutinho et al. (2025) Accelerated controller tuning using human feedback and Multi-Task Preferential Bayesian Optimization
  20. Coutinho et al. (2026) Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization
  21. Granley et al. (2023) Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses
  22. Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
  23. Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
  24. Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
  25. Mikkola et al. (2020) Projective Preferential Bayesian Optimization
  26. Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
  27. Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
  28. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  29. Menn et al. (2026a) Local Preferential Bayesian Optimization
  30. Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
  31. Zhu and Bemporad (2025) Global and Preference-Based Optimization with Mixed Variables Using Piecewise Affine Surrogates
  32. Theiner et al. (2026) Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models
  33. Haltia et al. (2026) Elicitation-Augmented Bayesian Optimization
  34. Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
  35. Ignatenko et al. (2025) On preference learning based on sequential Bayesian optimization with pairwise comparison
  36. Abdolshah et al. (2019) Multi-objective Bayesian optimisation with preferences over objectives
  37. Gardner et al. (2014) Bayesian Optimization with Inequality Constraints

28.8 Competing 'first' claims #

PBO sits between dueling bandits, preference-based reinforcement learning, control (the GLISp line), and human-computer interaction, and claims of being "first" in one community often overlook earlier work in another. Each claim below holds only within its exact setting (inference).

  • Constrained PBO (Iwai et al., 2025) claims to be the first to introduce inequality constraints, after C-GLISp (Zhu et al., 2022) handled unknown constraints in 2022 and a variant of StageOpt (Sui et al., 2018b) learned a utility from preference feedback under numerically measured safety constraints in 2018; it does not cite the validity labels of Benavoli et al. 2021 (Benavoli et al., 2021c). The claim holds for inequality constraints on a Gaussian process surrogate.
  • LineSpar (Cheng et al., 2020), a workshop paper, claims to be the first high-dimensional preference-based Bayesian optimization, after sequential line search (2017) and alongside projective PBO (ICML 2020).
  • DSTS (Astudillo et al., 2025) claims the first multi-objective PBO framework, which overlaps with a preprint on choice functions (Benavoli et al., 2021b); the claim holds for latent objectives observed only through preferences.
  • POP-BO (Xu et al., 2024b) says existing methods lack cumulative regret or global convergence guarantees, with the qualifier "continuous input space", and does not cite Kirschner and Krause (2021).
  • MPES (Nguyen et al., 2021) calls itself the first information-theoretic acquisition function with preference observations (first arXiv version December 2020), while the EIIG rule (Benavoli et al., 2021c) (August 2020) also uses the dueling information gain; both descend from BALD (Houlsby et al., 2011) (inference).
  • EUBO's one-step Bayes optimality was first shown in BOPE (Lin et al., 2022); the qEUBO paper's claim holds for its extension to logistic noise and q>2q > 2.
Sources cited in Section 28.8 12
  1. Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
  2. Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
  3. Sui et al. (2018b) Stagewise Safe Bayesian Optimization with Gaussian Processes
  4. Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
  5. Cheng et al. (2020) Preference-Based Bayesian Optimization in High Dimensions with Human Feedback
  6. Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
  7. Benavoli et al. (2021b) Choice functions based multi-objective Bayesian optimisation
  8. Xu et al. (2024b) Principled Preferential Bayesian Optimization
  9. Kirschner and Krause (2021) Bias-Robust Bayesian Optimization via Dueling Bandits
  10. Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
  11. Houlsby et al. (2011) Bayesian Active Learning for Classification and Preference Learning
  12. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes

28.9 The state of empirical comparison #

Table 28.3 lists the main comparison studies and their settings. In it, qTS is batch Thompson sampling, qNEI is noisy batch expected improvement, HPO-B is a benchmark built from hyperparameter optimization tasks, and a Sobol sequence is a quasi-random space-filling design.

Table 28.3 The main empirical comparisons of preferential acquisition rules and the settings in which they were run.
Study Compared Dimensions Budget and repetitions Noise model Main conclusion
González et al. 2017 (González et al., 2017) PE, CEI, DTS, random, interactive BO, Sparring 1, 2 5 + 200 duels; 20 repetitions not stated DTS best
Mikkola et al. 2020 (Mikkola et al., 2020) 5 projective rules against pairwise random and DTS variants 2, 6, 10, 20 100 queries; 25 initializations small Gaussian noise every projective variant beat every pairwise variant
Siivola et al. 2021 (Siivola et al., 2021) batch EI, batch Thompson sampling, random; three inference methods 6 functions; real data with d≤4d \le 4 batch sizes 2 to 6; 10 repetitions utility noise sd 0.05 differences between acquisition functions larger than between feedback types; only slightly better than baseline on real data
Nguyen et al. 2021 (Nguyen et al., 2021) MPES, EI, DTS 1 to 3; CIFAR-10 embedding; SUSHI not stated not stated MPES consistently best
Fauvel and Chalk 2021 (Fauvel and Chalk, 2021) 9 rules 34 functions 80 iterations; 40 repetitions probit, unit variance after normalization MUC and dueling UCB tied first; Brochu EI eighth; random ninth
Lin et al. 2022 (Lin et al., 2022) EUBO-ζ, EUBO-f̃, BALD-f̃, random, others multi-outcome problems 75 comparisons in 3 stages; 30 repetitions 10% wrong choices EUBO variants best
Astudillo et al. 2023 (Astudillo et al., 2023) qEUBO, MPES, qTS, qEI, qNEI, random 4 to 7 4d4d initial + 150 queries; 50 or 100 repetitions logistic, calibrated to 10%, 20%, 30% errors on the top 1% of point pairs with q=2q = 2, qEUBO best on all problems except Car cab
Takeno et al. 2023 (Takeno et al., 2023) HB, Laplace and EP with EI, MUC, skew-GP sampling rules up to 6 (12 functions) 3d3d initial; 10 repetitions noise variance 10−410^{-4} HB best overall; qEUBO not included
Xu et al. 2024, POP-BO (Xu et al., 2024b) DTS, HB, qEUBO GP samples; 6-D Ackley not stated logistic qEUBO's reported solution slightly better, cumulative regret more than 2.5 times higher
Zhang et al. 2025, PABBO (Zhang et al., 2025a) qEUBO, qEI, qNEI, qTS, MPES, random 1, 2, 6; HPO-B; Candy; Sushi 30 repetitions noise-free PABBO first or second; random often beat some GP baselines
Lazzaro et al. 2026, PF-TS (Lazzaro et al., 2026) MR-LPF, POP-BO, MaxMinLCB 1-D Ackley; 3-D catalyst T=300T = 300; 30 repetitions logistic PF-TS cumulative regret significantly lower than MR-LPF and POP-BO
Menn et al. 2026, local (preprint) (Menn et al., 2026a) local methods, qEUBO, HB with EI, GLISp, Sobol up to 102 about 10d10d comparisons, after 5d5d random evaluations for policy search Gaussian, 10% of the value range local methods better on steep optima

The noise calibration of the qEUBO experiments comes from their Appendix C.2 and code (Astudillo, 2023b).

The comparisons do not add up. The studies differ in metric (Section 31.4.3 lists five definitions of regret), in noise model, from a nearly noise-free probit to 10% flipped answers and moderate Gumbel noise (Section 31.4.2), and in posterior inference: Laplace, expectation propagation, variational inference, or skew Gaussian process sampling. Rankings flip with these choices, as POP-BO's comparison with qEUBO shows (inference). Fauvel and Chalk's Borda analysis over 34 functions is the broadest single comparison, but it predates qEUBO and the hallucination believer, uses expectation propagation throughout, and runs only 80 iterations. Takeno et al.'s comparison comes next, but its duels are almost noise-free and it does not include qEUBO. No comparison has breadth, realistic noise, and the acquisition functions of 2023 and later at the same time.

Random queries sometimes keep up. The PABBO authors write that the random strategy often beat some of the Gaussian process baselines (Zhang et al., 2025a). On real data with a low signal-to-noise ratio, Siivola et al. found that all methods only barely beat the baseline, and wrote that preference observations are inherently less informative than direct ones, that larger batches do not alleviate this, and that for very noisy data, random search with large batches may be a good choice (Siivola et al., 2021). In a single replication, BoTorch's BOPE tutorial prints candidate utilities of −0.473-0.473 for EUBO-ζ, −0.216-0.216 for random preference exploration, −0.101-0.101 for the true utility, and −1.380-1.380 for random experimentation, and notes that EUBO-ζ's win is not guaranteed in a single replication (Meta Platforms, Inc., 2026d). The three pieces of evidence point the same way: at small budgets and realistic noise, the gap between an elaborate acquisition function and a random design may be small, and a single run says little (inference).

Evidence from people. The human studies in this chapter are all small, and none was designed to compare acquisition functions: projective PBO's materials-science user, the Sequential Gallery's 6 participants, LineCoSpar's 6, constrained PBO's 11 designers, and Meta-PO's 36. As of September 2026, we found no study that randomizes people to different acquisition functions, such as qEUBO, DTS, and MUC, under the same interface and budget. Section 47.4 describes the experiment that would settle it. Until then, a practitioner should treat the ranking of acquisition functions as unknown for people and keep a share of random queries in every session as a control, which costs little given the evidence above (inference). Nor is there a benchmark for PBO with agreed metrics, noise models, and budgets: each paper reuses Forrester, six-hump camel, Hartmann, Ackley, Sushi, and Candy as it sees fit (Section 31.5).

Sources cited in Section 28.9 14
  1. González et al. (2017) Preferential Bayesian Optimization
  2. Mikkola et al. (2020) Projective Preferential Bayesian Optimization
  3. Siivola et al. (2021) Preferential Batch Bayesian Optimization
  4. Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
  5. Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
  6. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  7. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  8. Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
  9. Xu et al. (2024b) Principled Preferential Bayesian Optimization
  10. Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
  11. Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
  12. Menn et al. (2026a) Local Preferential Bayesian Optimization
  13. Astudillo (2023b) qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py)
  14. Meta Platforms, Inc. (2026d) Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1)

28.10 Settled, contested, missing #

Research status Settled, contested, missing

Settled. qEI is not asymptotically consistent on the instances the qEUBO paper constructs (Astudillo et al., 2023). Adapted expected improvement stalled in the experiments of four independent groups. EUBO's one-step Bayes optimality holds without noise, has an additive-constant guarantee under logistic noise, and was first shown in BOPE (Lin et al., 2022). Query forms in which each action carries more information, projections and planes, beat pairs in the simulations of the papers that proposed them (Mikkola et al., 2020; Koyama et al., 2020).

Contested. The value of more options per query: qEUBO found q=4q = 4 clearly better than q=2q = 2, Siivola et al. found only marginal gains, and their acquisition functions and noise differ. The advantage of the hallucination believer: it is best near noise-free answers, and stuck in local optima under logistic noise according to POP-BO. EUBO's collapse and ill-conditioning: two 2026 preprints, not yet peer reviewed or tested on human data. The ranking of Thompson-sampling rules, which flips with dimension and setting.

Missing. A broad comparison of acquisition functions on a shared benchmark with matched noise and budget. An experiment that randomizes people to acquisition functions. Acquisition functions derived for slider, plane, and projective queries. A stopping rule with guarantees for pairwise Gaussian process models. Regret bounds for DTS and HB, and continuous-domain guarantees for qEUBO under noise (Chapter 29).

For a choice that must be made now. The evidence supports including random queries and simple baselines among the comparisons, and reporting dimension, noise, and metric. qEUBO with PairwiseGP is the best-maintained default, but that recommendation comes from the software ecosystem and from within the PBO framework, not from a direct comparison with simpler methods (inference).

Sources cited in Section 28.10 4
  1. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  2. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  3. Mikkola et al. (2020) Projective Preferential Bayesian Optimization
  4. Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization

28.11 Exercises #

Exercise 28.1

Theorem 2 of the qEUBO paper bounds the loss from logistic noise of scale τ\tau by τ W((q−1)/e)\tau\, W((q - 1)/e), where WW is the inverse of w↦weww \mapsto w e^w. Compute W((q−1)/e)W((q - 1)/e) for q=2q = 2 and q=4q = 4 by solving wew=(q−1)/ew e^w = (q - 1)/e, and say what the bound implies about showing more options per query when answers are noisy.

Solution

For q=2q = 2, solve wew=1/e≈0.368w e^w = 1/e \approx 0.368: w=0.28w = 0.28 gives 0.28⋅e0.28≈0.370.28 \cdot e^{0.28} \approx 0.37, so W(1/e)≈0.28W(1/e) \approx 0.28. For q=4q = 4, solve wew=3/e≈1.10w e^w = 3/e \approx 1.10: w=0.60w = 0.60 gives 0.60⋅1.82≈1.090.60 \cdot 1.82 \approx 1.09, so W(3/e)≈0.60W(3/e) \approx 0.60. The guaranteed loss relative to the noise-free one-step optimum roughly doubles from q=2q = 2 to q=4q = 4, in units of the noise scale τ\tau, while the value of a noise-free answer from four options is at least that from two. The bound alone does not say whether more options help; it says the guarantee loosens slowly, which is consistent with qEUBO's experiments finding q=4q = 4 better than q=2q = 2 at moderate noise.

Exercise 28.2

A rough way to see why Thompson sampling over-explores as the dimension grows. Suppose the posterior is confident within distance r=0.1r = 0.1 of each of n=20n = 20 observed points in the unit cube [0,1]d[0, 1]^d, and uncertain elsewhere. Estimate the fraction of the cube that is confident for d=1d = 1 and d=6d = 6, using the volume of a dd-dimensional ball, πd/2rd/Γ(d/2+1)\pi^{d/2} r^d / \Gamma(d/2 + 1), and ignoring overlaps and edges. Where will the maximum of one posterior sample tend to fall?

Solution

For d=1d = 1 the "ball" is an interval of length 2r=0.22r = 0.2, so 20 points could cover the whole interval (the estimate 20×0.2=420 \times 0.2 = 4 exceeds 1; overlaps make the true coverage at most 1). For d=6d = 6, the ball volume is π3r6/3!≈31.0×10−6/6≈5.2×10−6\pi^3 r^6 / 3! \approx 31.0 \times 10^{-6} / 6 \approx 5.2 \times 10^{-6}, so 20 balls cover about 10−410^{-4} of the cube. Almost all of the six-dimensional cube is uncertain, so a posterior sample has many chances to be high somewhere nobody has looked, and its maximum tends to land there. That is exploration by construction, and in six dimensions it rarely returns to refine the best region. This picture is our illustration of the mechanism, not a result of the cited papers, which report the over-exploration without isolating its cause.

Exercise 28.3

Using Table 28.3, find a study that compares qEUBO with the hallucination believer. Under which noise model, at which dimensions, and with which metric? What would a comparison need in order to settle whether either rule is better at human noise levels?

Solution

Only the POP-BO paper (Xu et al., 2024b) includes both (with DTS), under logistic noise, on Gaussian process samples and the 6-dimensional Ackley function, with an unstated budget. It reports that the hallucination believer got stuck in local optima, and that qEUBO's reported solution was slightly better than POP-BO's but its cumulative regret more than 2.5 times higher. Takeno et al., where the hallucination believer did best, did not include qEUBO and used noise variance 10−410^{-4}. A settling comparison would fix a noise model calibrated to human comparisons, run both rules with the same surrogate, inference, budget, and dimensions, report both simple and cumulative regret, include random queries, and ideally be repeated with people.

Sources cited in Section 28.11 1
  1. Xu et al. (2024b) Principled Preferential Bayesian Optimization

Further reading #

  • Astudillo et al. (2023) is the reference for qEUBO; read its four theorems with their conditions, and its Appendix C.2 on how the noise was calibrated.
  • Fauvel and Chalk (2021) is the broadest comparison of rules (34 functions, nine rules) and the clearest treatment of epistemic against aleatoric uncertainty in acquisition.
  • Takeno et al. (2023) and Xu et al. (2024b) disagree about the hallucination believer; together they show how much the noise level decides.
  • Wu and Gardner (2026) derives the exact knowledge gradient under a probit likelihood and is the best place to see why EUBO's equivalence with it breaks under noise.
  • Mikkola et al. (2020) and Koyama et al. (2020) are the two reference papers on queries richer than a pair.

References

  1. Abdolshah, M., Shilton, A., Rana, S., Gupta, S., and Venkatesh, S. (2019). Multi-objective Bayesian optimisation with preferences over objectives. Advances in Neural Information Processing Systems. Cited in §28.7
  2. Adachi, M., Chau, S. L., Xu, W., Singh, A., Osborne, M. A., and Muandet, K. (2025). Bayesian Optimization for Building Social-Influence-Free Consensus. arXiv. preprint Cited in §28.6
  3. Astudillo, R. (2023b). qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py). GitHub. software Cited in §28.9
  4. Astudillo, R., and Frazier, P. (2020). Multi-attribute Bayesian optimization with interactive preference learning. International Conference on Artificial Intelligence and Statistics. Cited in §28.6 §28.7
  5. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §28.1 §28.2 §28.4 §28.5 §28.6 §28.7 §28.9 §28.10
  6. Astudillo, R., Li, K., Tucker, M., Cheng, C. X., Ames, A. D., and Yue, Y. (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Cited in §28.5 §28.7 §28.8
  7. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Cited in §28.6
  8. Benavoli, A., Azzimonti, D., and Piga, D. (2021b). Choice functions based multi-objective Bayesian optimisation. arXiv. preprint Cited in §28.8
  9. Benavoli, A., Azzimonti, D., and Piga, D. (2021c). Preferential Bayesian optimisation with skew gaussian processes. Proceedings of the Genetic and Evolutionary Computation Conference Companion. Cited in §28.1 §28.5 §28.7 §28.8
  10. Benavoli, A., Azzimonti, D., and Piga, D. (2023). Learning Choice Functions with Gaussian Processes. Uncertainty in Artificial Intelligence. Cited in §28.6
  11. Bıyık, E., Palan, M., Landolfi, N. C., Losey, D. P., and Sadigh, D. (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Cited in §28.1 §28.6 §28.7
  12. Brochu, E., de Freitas, N., and Ghosh, A. (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Cited in §28.1 §28.5
  13. Chen, E., Truong, S. T., Dullerud, N., Koyejo, S., and Guestrin, C. (2026). Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds. Conference on Uncertainty in Artificial Intelligence. Cited in §28.7
  14. Cheng, M., Novoseller, E., Tucker, M., Cheng, R., Yue, Y., and Burdick, J. (2020). Preference-Based Bayesian Optimization in High Dimensions with Human Feedback. SCMLS 2020 Workshop. workshop paper Cited in §28.8
  15. Cosner, R., Tucker, M., Taylor, A., Li, K., Molnár, T., Ubelacker, W., … Ames, A. (2022). Safety-Aware Preference-Based Learning for Safety-Critical Control. Learning for Dynamics and Control Conference. Cited in §28.7
  16. Coutinho, J. P. L., Peng, Y., Rendall, R., Rizzo, C., Ma, K., Chin, S.-T., Castillo, I., and Reis, M. S. (2025). Accelerated controller tuning using human feedback and Multi-Task Preferential Bayesian Optimization. 2025 American Control Conference (ACC). Cited in §28.7
  17. Coutinho, J. P., Peng, Y., Rendall, R., Ma, K., Chin, S.-T., Castillo, I., and Reis, M. S. (2026). Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization. Control Engineering Practice. Cited in §28.7
  18. Erarslan, A., Sevilla Salcedo, C., Tanskanen, V., Nisov, A., Päiväkumpu, E., Aisala, H., … Mikkola, P. (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Cited in §28.3
  19. Fauvel, T., and Chalk, M. (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. preprint Cited in §28.1 §28.4 §28.5 §28.9
  20. Gardner, J. R., Kusner, M. J., Xu, Z., Weinberger, K. Q., and Cunningham, J. P. (2014). Bayesian Optimization with Inequality Constraints. Proceedings of the 31st International Conference on Machine Learning (ICML 2014). Cited in §28.7
  21. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §28.1 §28.4 §28.5 §28.9
  22. Granley, J., Fauvel, T., Chalk, M., and Beyeler, M. (2023). Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. NeurIPS 2023. Cited in §28.7
  23. Haltia, A., Hyvönen, V., and Kaski, S. (2026). Elicitation-Augmented Bayesian Optimization. arXiv. preprint Cited in §28.3 §28.7
  24. Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M. (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Cited in §28.1 §28.5 §28.8
  25. Huber, F., Rojas Gonzalez, S., and Astudillo, R. (2025). Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization. Journal of Multi-Criteria Decision Analysis. Cited in §28.7
  26. Ignatenko, T., Kondrashov, K., Cox, M., and de Vries, B. (2025). On preference learning based on sequential Bayesian optimization with pairwise comparison. Artificial Intelligence. Cited in §28.2 §28.7
  27. Ip, J. H. S., Chakrabarty, A., Mesbah, A., and Romeres, D. (2025). User Preference Meets Pareto-Optimality in Multi-Objective Bayesian Optimization. Proceedings of the AAAI Conference on Artificial Intelligence. Cited in §28.7
  28. Iwai, K., Kumagae, Y., Koyama, Y., Hamasaki, M., and Goto, M. (2025). Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. Cited in §28.7 §28.8
  29. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §28.5
  30. Khan, F. A., Chakraborty, T., Dietrich, J. P., and Wirth, C. (2025). Efficient Contextual Preferential Bayesian Optimization with Historical Examples. Proceedings of the Genetic and Evolutionary Computation Conference Companion. Cited in §28.7
  31. Kirschner, J., and Krause, A. (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. Cited in §28.1 §28.8
  32. Kobalczyk, K., Lin, Z. J., Letham, B., Zhao, Z., Balandat, M., and Bakshy, E. (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Cited in §28.6
  33. Koyama, Y., Sato, I., Sakamoto, D., and Igarashi, T. (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Cited in §28.7
  34. Koyama, Y., Sato, I., and Goto, M. (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Cited in §28.6 §28.7 §28.10
  35. Kwon, Y., Tsurumine, Y., Shimmura, T., Kawamura, S., and Matsubara, T. (2022). Physically Consistent Preferential Bayesian Optimization for Food Arrangement. IEEE Robotics and Automation Letters. Cited in §28.7
  36. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §28.3 §28.5 §28.9
  37. Li, K., Tucker, M., Bıyık, E., Novoseller, E., Burdick, J. W., Sui, Y., … Ames, A. D. (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Cited in §28.6 §28.7
  38. Li, Z., Liao, Y.-C., and Holz, C. (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Cited in §28.7
  39. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §28.2 §28.4 §28.5 §28.7 §28.8 §28.9 §28.10
  40. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §28.7
  41. Menn, J., Kober, M., Brunzema, P., Stenger, D., and Trimpe, S. (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Cited in §28.3 §28.7 §28.9
  42. Menn, J., Stenger, D., and Trimpe, S. (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Cited in §28.7
  43. Meta Platforms, Inc. (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. software Cited in §28.2
  44. Meta Platforms, Inc. (2026d). Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1). botorch.org. software Cited in §28.6 §28.9
  45. Meta Platforms, Inc. (2026e). BoTorch CHANGELOG. GitHub. software Cited in §28.2 §28.5
  46. Mikkola, P., Todorović, M., Järvi, J., Rinke, P., and Kaski, S. (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §28.1 §28.6 §28.7 §28.9 §28.10
  47. Nguyen, Q. P., Tay, S., Low, B. K. H., and Jaillet, P. (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Cited in §28.1 §28.5 §28.8 §28.9
  48. Ozaki, R., Ishikawa, K., Kanzaki, Y., Takeno, S., Takeuchi, I., and Karasuyama, M. (2024). Multi-Objective Bayesian Optimization with Active Preference Learning. Proceedings of the AAAI Conference on Artificial Intelligence. Cited in §28.3 §28.6 §28.7
  49. Pásztor, B., Kassraie, P., and Krause, A. (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Cited in §28.3 §28.5
  50. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §28.3 §28.4 §28.5
  51. Siivola, E., Dhaka, A. K., Andersen, M. R., González, J., García Moreno, P., and Vehtari, A. (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Cited in §28.1 §28.5 §28.6 §28.9
  52. Sinaga, M. A., Martinelli, J., and Kaski, S. (2026). Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization. Symposium on Probabilistic Machine Learning (ProbML 2026), Proceedings Track. Cited in §28.3
  53. Sui, Y., Zhuang, V., Burdick, J. W., and Yue, Y. (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. Cited in §28.1 §28.5
  54. Sui, Y., Zhuang, V., Burdick, J., and Yue, Y. (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. Cited in §28.7 §28.8
  55. Takeno, S., Nomura, M., and Karasuyama, M. (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §28.2 §28.4 §28.5 §28.9
  56. Theiner, L., Pfefferkorn, M., Zhao, Y., Hirt, S., and Findeisen, R. (2026). Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models. European Control Conference. Cited in §28.7
  57. Tucker, M., Cheng, M., Novoseller, E., Cheng, R., Yue, Y., Burdick, J. W., and Ames, A. D. (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Cited in §28.6 §28.7
  58. Tucker, M., Novoseller, E., Kann, C., Sui, Y., Yue, Y., Burdick, J. W., and Ames, A. D. (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Cited in §28.6
  59. Wang, H., Branke, J., and Poloczek, M. (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. Cited in §28.7
  60. Wang, X., Zeng, Q., Zuo, J., Liu, X., Hajiesmaili, M., Lui, J. C., and Wierman, A. (2025b). Fusing Reward and Dueling Feedback in Stochastic Bandits. International Conference on Machine Learning. Cited in §28.6
  61. Wang, W., Shi, J., and Jones, C. N. (2025d). Personalized Building Climate Control with Contextual Preferential Bayesian Optimization. arXiv. preprint Cited in §28.7
  62. Wu, K., and Gardner, J. R. (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Cited in §28.3 §28.4 §28.5
  63. Xu, Y., Joshi, A., Singh, A., and Dubrawski, A. (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. Cited in §28.6
  64. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §28.3 §28.4 §28.5 §28.8 §28.9 §28.11
  65. Zhang, X., Huang, D., Kaski, S., and Martinelli, J. (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §28.3 §28.5 §28.7 §28.9
  66. Zhu, M., and Bemporad, A. (2025). Global and Preference-Based Optimization with Mixed Variables Using Piecewise Affine Surrogates. Journal of Optimization Theory and Applications. Cited in §28.7
  67. Zhu, M., Piga, D., and Bemporad, A. (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Cited in §28.7 §28.8