Software, Evaluation, and the Research Community
The previous chapters reported what methods for preferential Bayesian optimization (PBO) claim. This chapter turns to the conditions under which those claims were made and can be used: the software that implements the methods, the defaults it chooses on the user's behalf, the way papers evaluate one method against another, and the people who do the work.
The picture as of September 2026 has three parts. There is one maintained, general-purpose software stack for PBO, Meta's BoTorch and Ax; the rest is mostly code released with single papers. Evaluation relies almost entirely on simulated users answering comparisons about synthetic test functions inherited from scalar optimization. And the research community is small and spread over machine learning, control, robotics, and human-computer interaction, which publish in different venues and, in the cross-citations we checked, rarely cite one another.
31.1 Software #
A reader who wants to run PBO today has to choose software, and the choice is narrower than the size of the Bayesian optimization ecosystem suggests. Dates below are those in the projects' changelogs (the files in which a project lists what changed in each version) or, where a project has no changelog entry, upload dates on the Python Package Index (PyPI); the two can differ by a day.
31.1.1 BoTorch #
BoTorch is Meta's library of Gaussian process models and acquisition functions
built on PyTorch; Section 18.6 and Section 27.5 described its
preference model, PairwiseGP. Its preference features arrived between 2020 and
2024 (Meta Platforms, Inc., 2026e):
- 0.2.3 (April 27, 2020) added
PairwiseGPfor pairwise comparison data. - 0.3.2 (October 2020) removed its noise term and added a
ScaleKernelby default. - 0.6.3 (March 28, 2022) added
LearnedObjective, the analytic form of EUBO (the expected utility of the best option, Section 19.4), and a tutorial on Bayesian optimization with preference exploration (BOPE), in which a person states preferences over the outcomes of experiments rather than over designs. - 0.6.5 (July 2022) added a logistic link likelihood and switched the PBO tutorial to EUBO.
- 0.9.0 (August 1, 2023) added pairwise BALD (Bayesian active learning by disagreement, which picks the query whose answer the model's plausible hypotheses disagree about most).
- 0.10.0 (February 26, 2024) added qEUBO, the form of EUBO for queries of options.
Since then the preference code has received only fixes: in 0.18.0 (June 3,
2026), a fix to how PairwiseGP manages its stored state and a numerically
stabler mixture entropy in BALD. Version 0.18.1 (June 8, 2026) is the latest
release, and the last commit to the main branch was on September 30, 2026
(Meta Platforms, Inc., 2026i). The preference acquisition functions live in one source file
(Meta Platforms, Inc., 2026j).
BoTorch's PBO tutorial is the first code most newcomers run, so its design is worth knowing. It uses a 4-dimensional linear utility, adds Gaussian noise with standard deviation 0.1 to the utilities before comparing, evaluates the model with Kendall's rank correlation (the fraction of pairs that two rankings order the same way, rescaled to run from to ), and compares EUBO with random queries over only 3 repetitions (Meta Platforms, Inc., 2026c).
31.1.2 Ax #
Ax is Meta's platform for managing experiments, which calls BoTorch for its
models. A pairwise model bridge, PairwiseModelBridge, first appears in the
installed package of Ax 0.2.6 (August 17, 2022) (Facebook, Inc., 2022). Ax 1.0.0 was
uploaded to PyPI on May 8, 2025, with a companion paper by Olson et al. (2025) (AutoML
2025). Ax's preference features arrived in 2026 (Meta Platforms, Inc., 2026l):
- 1.2.2 (January 2026; PyPI has no files for this version) added a
PreferenceOptimizationConfigwith storage support and a Kendall rank correlation diagnostic. - 1.2.3 (February 19, 2026) added BOPE utility traces based on
PairwiseGP, andLLMProviderandLLMMessageabstractions for language models. - 1.3.0 (June 4, 2026) added LILO labeling trials with a check that the labels
are fresh, automatic dispatch of qEUBO for PBO,
PairwiseGPinsideModelList, and a utility ranking plot. LILO turns free-text feedback into pairwise labels with a language model (Section 35.2). - 1.3.1 (June 9, 2026) is the latest release and requires Python 3.11 or later (Meta Platforms, Inc., 2026a).
Ax's changelog begins only with 1.1.0 (August 2025), so the earlier history comes from the installed packages. The main branch (last commit September 29, 2026) has no tutorial on preferences, BOPE, or LILO (Meta Platforms, Inc., 2026m), and the LILO logic exists only in code (Meta Platforms, Inc., 2026b).
31.1.3 Every package with preference support #
Table 31.1 lists the software with preference support that we found; the links behind each name are code repositories or software documentation. HB-EI and HB-UCB are the hallucination believer of Section 19.3 with expected improvement or an upper confidence bound.
| Software | Surrogate and inference | Acquisition | Preference features | Latest release | Status |
|---|---|---|---|---|---|
| BoTorch (Meta Platforms, Inc., 2026f) | PairwiseGP: probit or logistic likelihood, Laplace approximation, hyperparameters by the Laplace marginal likelihood |
EUBO, qEUBO, pairwise BALD, pairwise posterior variance | models for pairwise data, learned objectives, PBO and BOPE tutorials | 0.18.1 (2026-06-08) | active; preference module only fixed since February 2024 |
| Ax | calls BoTorch's PairwiseGP |
qEUBO dispatched automatically (from 1.3.0) | preference optimization config, BOPE utility traces, LILO labeling trials, utility ranking plot, Kendall rank correlation diagnostic | 1.3.1 (2026-06-09) | active; no preference tutorial |
| AEPsych (Meta, 2026) | PairwiseProbitModel, inheriting from PairwiseGP |
we did not check | pairwise psychophysics experiments | 0.8.0 (2025-04-11); main branch 0.8.0+dev | commits on 2026-08-26 (UTC) |
| optuna-dashboard (Optuna developers, 2026b) | after Takeno et al. 2023: expectation propagation to fit kernel hyperparameters, Gibbs sampling of the latent utility | log expected improvement maximized on a sampled Gaussian process | web interface; the user marks the worst of several candidates (4 in the tutorial) and every implied pair is recorded; available since 0.13.0b1 (September 2023) | 0.21.0 (2026-09-10) | active; no dynamic search spaces |
| OptunaHub plmbo (Ozaki, 2026) | GPy Gaussian processes with an RBF kernel for each objective; preference weights by NumPyro Markov chain Monte Carlo | we did not check | implements the multi-objective BO with active preference learning of Ozaki et al. (2024) (AAAI 2024), with comparisons typed at a console | optunahub 0.5.0 (2026-09-09); plugin declares Optuna 3.6.1 | OptunaHub has no purely preferential sampler |
| pySequentialLineSearch (Koyama, 2025b) (C++ with Eigen and NLopt; experimental Python bindings) | PreferenceRegressor: Bradley-Terry-Luce likelihood, Matérn 5/2 kernel, maximum a posteriori hyperparameters |
expected improvement (default), GP upper confidence bound | slider-based line search; pairwise comparison demo | tag v0.5 (2025-11-11), setup.py still says 0.4; not on PyPI |
maintained at low frequency; tested only on macOS 10.15 |
| GLIS, GLISp, C-GLISp (Bemporad, 2023) (Python and MATLAB) | radial basis function surrogate (inverse quadratic by default) with inverse distance weighting; not probabilistic | surrogate minus an inverse-distance-weighting variance term and an exploration term, solved by particle swarm optimization | preferences in (ties allowed); C-GLISp adds unknown constraints and a "satisfactory" label | glis 2.0.2 (2023-06-19) | no commits after 2023-03-03; mixed-variable successor PWAS/PWASp (Zhu, 2025), last commit 2025-01-23 |
| prefGP (Benavoli and Azzimonti, 2026b) (JAX and PyTorch) | 9 preference and choice likelihoods; Laplace, variational inference (full and sparse), slice sampling | none | preference learning only, no optimization loop | no release tag; not on PyPI | last commit 2026-07-14; pins botorch 0.9.3, gpytorch 1.11, jax 0.4.23, torch 2.1.2 |
| SkewGP (Benavoli et al., 2025) | closed-form skew Gaussian process posterior, optional Laplace | Thompson sampling, upper confidence bound, an information-gain rule | preference and mixed data | no tag | last commit 2025-08-05 (README only) |
| preferentialBO (CyberAgent AI Lab, 2023) (Python 3.9, GPy 1.10.0) | Gibbs-sampled preference Gaussian process; expectation propagation, Laplace, and MCMC baselines | HB-EI, HB-UCB, and the paper's 6 baseline rules | reproduction code for Takeno et al. (ICML 2023) | no tag | no commits after 2023-07-26 |
| qEUBO paper code (Meta Research, 2023; Astudillo, 2023a) | in the paper, a variational Gaussian process with inducing points | qEUBO and baselines | the official repository has 5 tasks; the author's repository has more, including the Animation data | no tag | no commits after March 2023; no dependency file |
| PABBO (Zhang, 2026) | transformer policy pretrained by reinforcement learning | the policy outputs the query pair | amortized preference optimization | no tag | last commit 2026-09-14; pins botorch 0.10.0 and a CUDA 11.8 build of torch; needs a Weights and Biases key |
| Emukit example (Emukit developers, 2026) (GPy, Stan) | expectation propagation, variational inference, MCMC | batch comparison rules | batch PBO, in the examples directory only | emukit 0.5.1 (2026-02-22) | example unmaintained; still requires Python 3.7, scipy 1.1.0, and pystan |
Several more methods exist only as code for a single paper: PPBO (Aalto PML, 2022) (MIT, last commit September 13, 2022); POLAR (Tucker, 2024) (MATLAB, BSD-3-Clause, last commit June 17, 2024, supporting pairwise, coactive, and ordinal feedback); the LILO research code (Meta Research, 2026) (MIT, last commit May 12, 2026); and DT-PBO, PBO with decision-tree surrogates (Quadt, 2026) (MIT). Three repositories have no license file, so default copyright applies and reuse is not permitted without asking: POP-BO (PREDICT-EPFL, 2024), MR-LPF (Kayal, 2025), and CrashPBO (Institute for Data Science in Mechanical Engineering, RWTH Aachen University, 2026).
31.1.4 Libraries without preference support #
None of the major Bayesian optimization and Gaussian process libraries has a
preference likelihood or a preference acquisition function. We searched the
source trees of the latest versions at the end of September 2026, GPyTorch
1.15.2 (GPyTorch developers, 2026), GPflow 2.11.1 (GPflow developers, 2026), Trieste 4.6.0
(Secondmind Labs, 2026), SMAC 2.4.1 (AutoML.org, 2026), HEBO 0.3.6
(Huawei Noah's Ark Lab, 2024), Dragonfly 0.1.7 (Dragonfly developers, 2022), GPyOpt 1.2.6
(SheffieldML, 2023), and Optuna 5.0.0 (Optuna developers, 2026a), for pairwise,
preference, and dueling. PairwiseGP uses GPyTorch's kernels but
implements its own Laplace inference; "dueling" appears in HEBO only in an
unrelated reinforcement learning subproject; and Optuna's preference support
lives in optuna-dashboard and OptunaHub. GPyOpt was archived on January 17,
2023, and Dragonfly is effectively unmaintained.
31.1.5 What the software landscape means #
Three judgments follow (inference). Practical PBO depends
on one organization: BoTorch supplies the models and acquisition functions, Ax
manages experiments, BOPE, and LILO, and AEPsych serves psychophysics; since
qEUBO in February 2024, new preference features have landed in Ax rather than
BoTorch. Alternatives to the Gaussian process (GLISp, DT-PBO, PABBO) are niche,
and only GLIS can be installed with pip. And for a practitioner, Ax 1.3.x is
the most complete maintained software but lacks documentation for these
features; we did not verify whether Ax lets a user pass a dimension-scaled
prior to PairwiseGP through its model configuration.
Sources cited in Section 31.1 40
- Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
- Meta Platforms, Inc. (2026i) botorch release history
- Meta Platforms, Inc. (2026j) botorch/acquisition/preference.py
- Meta Platforms, Inc. (2026c) Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1)
- Facebook, Inc. (2022) ax-platform 0.2.6
- Olson et al. (2025) Ax: A Platform for Adaptive Experimentation
- Meta Platforms, Inc. (2026l) CHANGELOG (versions 1.2 to 1.3)
- Meta Platforms, Inc. (2026a) ax-platform release history
- Meta Platforms, Inc. (2026m) tutorials directory
- Meta Platforms, Inc. (2026b) ax/generation_strategy/transition_criterion.py
- Meta Platforms, Inc. (2026f) BoTorch LICENSE
- Meta (2026) AEPsych
- Optuna developers (2026b) optuna-dashboard 0.21.0
- Ozaki (2026) PLMBO (Preference Learning Multi-Objective Bayesian Optimization)
- Ozaki et al. (2024) Multi-Objective Bayesian Optimization with Active Preference Learning
- Koyama (2025b) sequential-line-search
- Bemporad (2023) GLIS
- Zhu (2025) PWAS
- Benavoli and Azzimonti (2026b) prefGP
- Benavoli et al. (2025) SkewGP
- CyberAgent AI Lab (2023) preferentialBO
- Meta Research (2023) qEUBO
- Astudillo (2023a) qEUBO
- Zhang (2026) PABBO
- Emukit developers (2026) preferential_batch_bayesian_optimization example
- Aalto PML (2022) PPBO
- Tucker (2024) POLAR
- Meta Research (2026) lilo
- Quadt (2026) DT-PBO-preprint
- PREDICT-EPFL (2024) POP-BO
- Kayal (2025) BOHF_code_submission
- Institute for Data Science in Mechanical Engineering, RWTH Aachen University (2026) crashpbo
- GPyTorch developers (2026) gpytorch 1.15.2
- GPflow developers (2026) gpflow 2.11.1
- Secondmind Labs (2026) trieste 4.6.0
- AutoML.org (2026) smac 2.4.1
- Huawei Noah's Ark Lab (2024) HEBO 0.3.6
- Dragonfly developers (2022) dragonfly-opt 0.1.7
- SheffieldML (2023) GPyOpt (archived)
- Optuna developers (2026a) optuna 5.0.0
31.2 Default priors that ignore dimension #
A default prior is a modeling decision that most users never see.
Section 27.5 read the choices inside PairwiseGP (Table 18.2
lists them), and Section 30.1 explained why a lengthscale prior of fixed
scale fails as the number of inputs grows. Table 31.2 collects the
defaults of every preference package we examined, with BoTorch's scalar models
as the point of comparison. No row scales with
dimension except the comparison row.
| Software or paper | Kernel | Default lengthscale setting | Mode or initial value | Scales with dimension |
|---|---|---|---|---|
BoTorch PairwiseGP (Meta Platforms, Inc., 2026h) |
RBF, inside a required ScaleKernel |
Gamma(2.4, 2.7), initialized at its mode, lower bound | about 0.52 | no |
| BoTorch's other models from 0.12.0 (comparison) (Meta Platforms, Inc., 2026k) | depends on the model | log-normal with location and scale , lower bound 0.025 | about 0.65 at , about 2.05 at (inference) | yes |
| optuna-dashboard preferential sampler (Optuna developers, 2026c) | Matérn 3/2, one lengthscale per input | Gamma(5, 10); noise prior Gamma(5, 50) | 0.4 | no |
| sequential-line-search (Koyama, 2025a) | Matérn 5/2, one lengthscale per input | default 0.5, maximum a posteriori prior with variance 0.25 | 0.5 | no |
| Koyama et al. 2017 and 2020 papers (Koyama et al., 2017; Koyama et al., 2020) | Matérn 5/2 in the 2020 paper | log-normal with parameters and , nearly fixed | about 0.5 | no |
| PABBO synthetic prior, in training and evaluation (Zhang et al., 2025a; Zhang, 2025) | RBF and Matérn 5/2, 3/2, 1/2 | truncated to , centered near 1/3 | not applicable | no |
No PBO package documents how its default prior relates to dimension, so anyone using these packages above about 10 dimensions inherits the fixed-scale priors that scalar Bayesian optimization abandoned in 2024 (inference; Figure 30.1 shows the consequence, and Section 30.3.5 what to pass instead). This is a different question from whether a prior should be informed by other users' data, a population prior, which Section 32.5 takes up: the defaults in Table 31.2 are weak priors, but weak at the wrong scale once the dimension grows (inference).
Sources cited in Section 31.2 8
- Meta Platforms, Inc. (2026h) BoTorch PairwiseGP source code pairwise_gp.py
- Meta Platforms, Inc. (2026k) botorch/models/utils/gpytorch_modules.py
- Optuna developers (2026c) optuna-dashboard PreferentialGPSampler source code gp.py
- Koyama (2025a) preference-regressor.hpp
- Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Zhang (2025) PABBO code repository: evaluation config evaluate.yaml
31.3 Reproducing research code #
A method that cannot be rerun cannot be checked, and a baseline that cannot be rerun gets reimplemented, each time a little differently. Three habits make research code rerunnable: a release tag (a named, frozen version of the repository), a dependency file that lists the exact library versions the code needs, and a license that says who may reuse it.
At the end of September 2026 we checked 11 research repositories, prefGP,
SkewGP, preferentialBO, the official qEUBO repository, PABBO, PPBO, POLAR, GLIS,
POP-BO, LILO, and BO_toolbox, for release tags: all 11 had none.
Dependencies are commonly pinned to versions from 2020 to 2024
(Table 31.1 lists several), PPBO pins numpy 1.18.4 and the archived
GPyOpt 1.2.6 (Aalto PML, 2022), and neither qEUBO repository has a dependency file
(Meta Research, 2023; Astudillo, 2023a). preferentialBO's scripts call
np.int, which NumPy 1.24 removed, so they run only with the pinned numpy
1.23.2; its README also warns that on Ubuntu 20.04 it has not been confirmed
that the paper's results can be fully reproduced (the experiments were run on
CentOS 6.9) (CyberAgent AI Lab, 2023).
Licenses are not uniform either. Most repositories use MIT, BSD, or Apache licenses, but PABBO uses the AGPL-3.0, a copyleft license that requires derived software, including software offered over a network, to be released under the same license; AEPsych's Creative Commons Attribution-NonCommercial 4.0 license restricts commercial reuse (Meta, 2026); and three repositories from 2024 to 2026 have no license file at all. Each method also brings its own Gaussian process implementation (GPy, BoTorch 0.9 or 0.10, JAX, MATLAB), so most papers reimplement their baselines, which may be one reason the same baseline looks strong in one paper and weak in another (inference).
Expect to pin old versions. A repository without a tag is a moving target, and one without a dependency file leaves the versions to guesswork: record the commit hash and the versions you got working. Check the license before building on the code; three of the repositories above have none, and one is copyleft. When a paper's baseline is a reimplementation, the baseline's settings are part of the result (inference).
Sources cited in Section 31.3 5
- Aalto PML (2022) PPBO
- Meta Research (2023) qEUBO
- Astudillo (2023a) qEUBO
- CyberAgent AI Lab (2023) preferentialBO
- Meta (2026) AEPsych
31.4 How methods are evaluated #
Comparing optimizers is harder than it looks. A single run depends on the random initial design and on the noise in the answers, so a method can win one run by luck, which is why comparisons average over many runs (Chapter 11). And a regret bound (Chapter 13) holds for a class of functions, not for the function in front of you, so only benchmarks can say how methods rank in practice, within the limits this section describes.
31.4.1 Inherited test functions and simulated users #
Since González et al. (2017), almost every PBO paper has been evaluated with a simulated user: a program that answers each comparison by evaluating a known function at the two options and adding noise. The functions are standard synthetic test functions, usually in 1 to 8 dimensions, which González et al. took from the library of scalar test functions maintained by Surjanovic and Bingham and later papers kept. These functions were designed to test scalar optimizers; they have none of the properties of human utility that Part IV and Part IX describe, such as indifference thresholds, drift, intransitivity, and anchoring (inference). Table 31.3 lists the test problems and metrics of representative papers; Table 28.3 gives the noise, budget, and conclusion of most of them, and the lists below collect the noise models and metrics.
| Paper | Test problems (dimension) | Main metric |
|---|---|---|
| González et al. (2017), ICML 2017 | Forrester (1); Six-hump camel, Gold-Stein, Levy (2); 33-point grid per dimension; 5 initial duels, total budget of 200 duels | true value at the current Condorcet winner |
| Mikkola et al. (2020), ICML 2020 | Six-hump camel (2), Hartmann (6), Levy (10), Ackley (20); 100 queries | true objective value at the maximizer of the posterior mean |
| Siivola et al. (2021), MLSP 2021 | several functions from the SigOpt library; 4 real data sets including Sushi and Candy () | we did not extract it |
| Fauvel and Chalk (2021), arXiv 2021 (preprint) | 34 functions from the Surjanovic and Bingham library, standardized to mean 0 and variance 1 | final optimum; Mann-Whitney U tests with Borda scores |
| BOPE: Lin et al. (2022), AISTATS 2022 | Vehicle safety (, outcomes), DTLZ2, OSY (, ), Car cab (, ), with several utilities | true utility at the maximizer of the posterior mean |
| qEUBO: Astudillo et al. (2023), AISTATS 2023 | Ackley (6), Alpine1 (7), Hartmann (6), Car cab (7), Sushi (4), Animation (5); initial queries, then 150 | log simple regret at the maximizer of the posterior mean |
| Takeno et al. (2023), ICML 2023 | 12 functions (8 in the main text, up to Hartmann in 6 dimensions) | regret at the recommended point |
| POP-BO: Xu et al. (2024b), ICML 2024 | Gaussian process samples, standard test functions including 6-dimensional Ackley, a thermal comfort problem | cumulative regret and suboptimality of the reported solution |
| MaxMinLCB: Pásztor et al. (2024), NeurIPS 2024 | Ackley, Eggholder, and others; Yelp restaurant data | cumulative regret in preference probabilities |
| PABBO: Zhang et al. (2025a), ICLR 2025 | Gaussian process samples (1, 2), Forrester, Beale, Branin; Ackley (6) and Hartmann (6) in an appendix; HPO-B, Candy, Sushi | simple regret of the best queried point |
| MR-LPF: Kayal et al. (2025), ICML 2025 | samples from a reproducing kernel Hilbert space, Ackley (1), Yelp (275 restaurants, 20 users) | cumulative regret in preference probabilities |
| PF-TS: Lazzaro et al. (2026), AISTATS 2026 | Ackley (1) with ; hydrogen yield of 63 catalyst compositions of three metals, with | cumulative regret in preference probabilities |
In the table, MaxMinLCB is the max-min lower confidence bound algorithm of Pásztor et al., PF-TS is Thompson sampling with preference feedback, and HPO-B is a benchmark built from hyperparameter tuning tasks of the kind Chapter 22 works through. Papers from 2026 continue the same practice, among them KappaSharp, a preprint, with 11 benchmarks in 5 to 20 dimensions (Shao et al., 2026).
31.4.2 Noise models chosen paper by paper #
Each paper picks its own model of how the simulated user errs. We found at least six kinds:
- Gaussian noise added to the utilities before comparing, with standard deviation 0.01 (Takeno et al.), 0.05 (Siivola et al.), 0.1 (the BoTorch tutorial), or 10% of the function's range (local PBO (Menn et al., 2026a)).
- Logistic noise calibrated so that random pairs among the top 1% of points are answered wrongly 10%, 20%, or 30% of the time (qEUBO, in an appendix study); for 6-dimensional Ackley, these correspond to logistic scales of 0.0575, 0.1416, and 0.2943 (Astudillo, 2023b).
- A fixed probability of flipping the answer (10% in BOPE).
- No noise by default (PABBO).
- Probit noise on a standardized function (Fauvel and Chalk).
- A decision maker simulated by a language model (LILO (Kobalczyk et al., 2026)).
Calibrating to an error rate is closer to reality than setting a scale directly, because it fixes how often the simulated user is wrong on the comparisons that matter, but the noise is still homoscedastic (the same everywhere) and independent from one answer to the next (inference). Even two models matched to the same error rate on one gap disagree on every other gap, as Figure 31.1 shows.
Some things to try:
- With the defaults, read the three curves at a gap of 0.30. The flip model still errs 10% of the time, logistic noise 0.14%, and probit noise 0.006%: "10% noise" makes a clearly worse option about 70 times more likely to win under the flip model than under logistic noise, and over a thousand times more likely than under probit noise (Exercise 31.3 works the probit case by hand).
- Move the reading gap to 0.05, a closer call than the matched one. Now logistic and probit noise err 25% and 26% of the time, the flip model still 10%: the flip model is harsher on clear choices and gentler on close ones.
- Raise the error rate at the matched gap to 30%. At a gap of 0.30 the logistic and probit users now err 7.3% and 5.8% of the time, close to each other, while the flip model errs 30%. With very noisy answers the two utility-noise models nearly agree; it is the flip model that stands apart.
31.4.3 Regret, defined five ways #
Regret is the shortfall of a method's result from the best achievable value (Section 13.1). In PBO the method never sees utility values, so a paper must decide which point counts as "the result", and the papers decide differently:
- the function value at the current Condorcet winner, the point that wins most often (González et al.);
- the simple regret at the maximizer of the posterior mean, the point the model would recommend (qEUBO, BOPE; Takeno et al. use their recommended point);
- the simple regret of the best point queried so far (PABBO);
- the cumulative regret summed over both points of every comparison, in preference probabilities (MaxMinLCB, MR-LPF, PF-TS);
- the cumulative regret in utility (POP-BO).
Papers also report the quality of the model's ranking (Kendall's rank correlation in the BoTorch tutorial and in Ax), the computing time per iteration (qEUBO, PABBO, Takeno et al.), and, in a user experiment, the number of simulation steps needed to reach the nearest local minimum (Mikkola et al.). No paper uses the number of comparisons needed to reach a threshold as its main metric, although that is the quantity a person in the loop pays for.
The choice of metric changes the ranking of methods. In the POP-BO paper, qEUBO reports a slightly better solution than POP-BO on Gaussian process sample instances, but its cumulative regret is more than 2.5 times higher, and on 6-dimensional Ackley the two metrics reverse again (Xu et al., 2024b). A method that queries near the optimum is favored by regret at the queried points; a method whose model is well calibrated is favored by regret at the posterior mean (inference). Figure 31.2 lets the reader watch this happen.
Some things to try:
- With the defaults (30 runs, noise 0.1), Recommended puts EUBO first, the incumbent rule close behind with overlapping bands, and random pairs last.
- Switch to Best queried. Random pairs now win, with regret near 0.05, through coverage rather than learning: 16 random pairs make 32 draws from only 41 candidates, so on average about 23 of the candidates have been shown (Exercise 31.1). In a larger space, or in more dimensions, the same metric would not favor random pairs so strongly (inference).
- Switch to Cumulative. The incumbent rule wins clearly: every pair it asks contains the model's current best guess, which is soon near the optimum, so half of each comparison costs almost nothing (Exercise 31.2).
- Set Runs to 5 and press Another batch of runs several times. Under Recommended, the winner changes from batch to batch, mostly between EUBO and the incumbent rule: with few repetitions, the ranking is partly a draw.
- Change the noise. At 0.2, random pairs overtake the incumbent rule under Recommended, while the other two metrics keep their order.
31.4.4 Statistics and the size of the differences #
Statistical reporting is thin. Papers use from 3 repetitions (the BoTorch tutorial), 10 (Takeno et al.), 20, and 30, up to 50 to 100 (qEUBO), and show uncertainty as ±1, ±1.96, or ±2 standard errors, 95% confidence intervals, or interquartile ranges. Only Fauvel and Chalk test many hypotheses formally: across 34 functions they compare every pair of methods with a Mann-Whitney U test (a test of whether one method's results tend to be larger than another's, without assuming a distribution) and aggregate the wins into Borda scores (each method scores one point for every method it beats) (Fauvel and Chalk, 2021). No later paper adopted the approach. Where differences are measured at all, the gap between elaborate acquisition functions and random queries is often small, and Section 28.9 reports the evidence and why the existing comparisons do not add up. The practical reading: repeat runs at least tens of times, include random queries, and test the differences that a claim rests on (inference).
Sources cited in Section 31.4 16
- González et al. (2017) Preferential Bayesian Optimization
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
- Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Menn et al. (2026a) Local Preferential Bayesian Optimization
- Astudillo (2023b) qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py)
- Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback
31.5 Real-data tasks and missing datasets #
The tasks that papers call "real data" all turn the judgments of a group into a single deterministic utility. Table 31.4 lists the four in use.
| Task | Source | What it contains | How PBO papers use it |
|---|---|---|---|
| Sushi | Kamishima's sushi preference data (Kamishima, 2026) | human rankings of sushi types | Siivola et al.: complete rankings of 100 sushi types with 4 continuous features; PABBO: five-point ratings averaged over users |
| Candy | FiveThirtyEight's online pairwise voting (FiveThirtyEight, 2017) | 2 features, sugar percentile and price percentile | both papers turn the votes into one complete ranking or win rate; Siivola et al. count 86 candies, PABBO 85 |
| Yelp | restaurant ratings | 275 restaurants, 20 users, 32-dimensional embeddings | MR-LPF and MaxMinLCB; built from ratings, not comparisons |
| Animation | qEUBO (Astudillo et al., 2023) | a fire-like particle effect with 5 parameters, from an AEPsych demo | the authors "collected 100 such pairwise comparisons from human users", fitted a model, and used it as the ground-truth test function |
These tasks reward the method that finds the optimum of a population, and they hide the inconsistency of individual judgments, which is the very thing PBO exists to handle (inference). We found no shared benchmark suite or leaderboard for PBO, and no public data set of individual-level pairwise judgments designed for evaluating PBO. Such a data set, with timestamps, presentation order, response times, and repeated pairs, is what would let methods be compared offline on real answers (Section 46.7).
Sources cited in Section 31.5 3
- Kamishima (2026) SUSHI Preference Data Sets
- FiveThirtyEight (2017) candy-power-ranking data
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
31.6 Reproduction problems found #
Checking papers against their code turned up the following problems.
- qEUBO's Animation task. The paper says that the fitted model is a support
vector machine (Astudillo et al., 2023). The author's personal repository has
two versions of the ground truth:
animation_runner.pyuses a support vector machine classifier andanimation2_runner.pyloads a fittedPairwiseGP(Astudillo, 2023a). The official repository the paper cites has no run script or data for the Animation task (Meta Research, 2023), and neither repository lists its dependencies. - BOPE's dimensions. The text gives "DTLZ2 (d = 4, k = 8)" while a figure is labeled "DTLZ2 (d=8, k=4)" (Lin et al., 2022).
- Candy's size. Two papers report 86 and 85 candies (Siivola et al., 2021; Zhang et al., 2025a); FiveThirtyEight's data file has 85 rows (FiveThirtyEight, 2017).
- Takeno et al.'s code does not run as is in a current environment (Section 31.3) (CyberAgent AI Lab, 2023).
None of these is large on its own; together they mean that the benchmark results of PBO cannot at present be reproduced one by one without contacting the authors (inference).
Sources cited in Section 31.6 8
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Astudillo (2023a) qEUBO
- Meta Research (2023) qEUBO
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- FiveThirtyEight (2017) candy-power-ranking data
- CyberAgent AI Lab (2023) preferentialBO
31.7 Simulated and real users #
Every benchmark above assumes that a simulated user stands in for a person well enough to rank methods. The few studies that compared the two directly all found clear differences. Table 31.5 collects them.
| Study | Setting | Participants | What differed from the simulated user |
|---|---|---|---|
| Schoinas et al. (2025), EMBC 2025 | retinal implant encoders, simulated prosthetic vision | 17 sighted | same choice as the simulated agent in only about 50% of trials; final loss 0.27 against 0.07 in simulation |
| Ou et al. (2022), Mensch und Computer 2022 | 9-parameter polygon simplification for 3D models | 2 artists for 3 months; 20 in the lab | satisfactory results in 11.9% of sequences in the field, 48.5% in the lab; inconsistent judgments, loss aversion |
| Ou et al. (2023), IUI 2023 | text, photo, and 3D mesh tasks | 60 | stopping depends on expertise |
| Colella et al. (2020), UMAP 2020 | 1-dimensional function, scalar feedback | 21 | users who understood the optimizer gave strategically biased answers |
| Chan et al. (2022), CHI 2022 | 3D touch interaction design, multi-objective BO on measured completion time and spatial error | 40 novice designers | better results but lower sense of agency and expressiveness |
| Taddei et al. (2026), arXiv 2026 (preprint) | preference optimization of an active prosthesis | simulations, then trials with 4 people (3 with one method, 1 with another) | in one of three trials with one participant, the preference estimate kept fluctuating after iteration 15, possibly from fatigue |
The table's numbers carry the main point, and three details add to it. In the retinal implant study, which tested an optimization that Granley et al. (2023) (NeurIPS 2023) had evaluated only in simulation, 16 of the 17 participants still preferred the optimized result under the main condition, and the authors stress "the importance of validating optimization strategies with human participants" (Schoinas et al., 2025). In Ou et al.'s three-month field deployment, 415 of 549 evaluation sequences stopped at the first iteration, and the authors conclude that "optimization using preferential choices lacks mechanisms to deal with inconsistent and contradictory human judgments", and that "machine outcomes, in turn, influence future user inputs via heuristic biases and loss aversion" (Ou et al., 2022) (Section 32.3, Section 32.6). And users who understand how the optimizer works "strategically provide biased answers" (Colella et al., 2020), contrary to the simulated user's assumption that feedback is faithful, while Mikkola et al. report that their method could tell choices made by people from those made by a computer program (Mikkola et al., 2020).
None of these behaviors appears in any simulated-user benchmark we found (inference). Simulated evaluation can compare how algorithms behave under a given noise model, but it cannot predict how they rank with people (Section 28.9 describes the missing experiment). Until simulators are validated on individual data, a result obtained with a simulated user should be checked on at least a few real people before it is reported as a result about people (inference); the norms that experimental economics adopted for such comparisons are a useful model (Section 40.12).
Sources cited in Section 31.7 8
- Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Ou et al. (2023) The Impact of Expertise in the Loop for Exploring Machine Rationality
- Colella et al. (2020) Human Strategic Steering Improves Performance of Interactive Optimization
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Taddei et al. (2026) Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis
- Granley et al. (2023) Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
31.8 The research community #
PBO is the work of a small number of groups. This section maps them by discipline, then reports the infrastructure of a field (surveys, theses, workshops, venues) and how much is published.
31.8.1 Groups by discipline #
Table 31.6 arranges the groups by the discipline they publish in, which is also, largely, the kind of question they ask.
| Discipline (typical venues) | Group (institutions) | Representative work | Direction |
|---|---|---|---|
| Machine learning methods (ICML, AISTATS, NeurIPS, ICLR, TMLR) | Frazier, Astudillo, Bakshy, Lin (Cornell University, Meta, Caltech) | (Astudillo and Frazier, 2020); BOPE (Lin et al., 2022); qEUBO (Astudillo et al., 2023); (Astudillo et al., 2025); LILO (Kobalczyk et al., 2026) | decision-theoretic acquisition and preference exploration over outcomes, recently labels from language models |
| Takeno, Nomura, Karasuyama (Nagoya Institute of Technology, CyberAgent AI Lab, RIKEN) | (Takeno et al., 2023; Ozaki et al., 2024) | skew Gaussian process inference and the hallucination believer; the optuna-dashboard sampler implements their method | |
| Benavoli, Azzimonti, Piga (Trinity College Dublin, IDSIA) | (Benavoli et al., 2021c; Benavoli and Azzimonti, 2026a) | preference and choice likelihoods, exact posteriors, the prefGP code | |
| Kaski (Aalto University) | PPBO (Mikkola et al., 2020); batch PBO (Siivola et al., 2021); PABBO (Zhang et al., 2025a); (Sinaga et al., 2026) | projective queries, batches, amortization, noise models | |
| Machine learning theory (ICML, NeurIPS, AISTATS) | Krause (ETH Zurich) | (Kirschner and Krause, 2021); MaxMinLCB (Pásztor et al., 2024) | dueling-bandit theory |
| Vakili and colleagues (MediaTek Research; collaborators at University College London and Imperial College) | MR-LPF (Kayal et al., 2025); PF-TS (Lazzaro et al., 2026) | regret theory; no software releases, no human studies | |
| Jones (EPFL) | POP-BO (Xu et al., 2024b) | guarantees, with building thermal comfort as the application | |
| Control (ECC, ACC, IEEE Transactions on Control Systems Technology) | Bemporad, Piga (IMT School for Advanced Studies Lucca; IDSIA) | GLISp (Bemporad and Piga, 2021); C-GLISp (Zhu et al., 2022); PWAS and PWASp (code (Zhu, 2025)) | non-probabilistic surrogates and controller calibration |
| Robotics (ICRA, IROS, IEEE Robotics and Automation Letters) | Ames, Yue, Tucker, Novoseller (Caltech) | CoSpar (Tucker et al., 2020b); LineCoSpar (Tucker et al., 2020a); POLAR (Tucker et al., 2022) | personalizing exoskeleton gaits (Chapter 24) |
| Trimpe (RWTH Aachen University) | CrashPBO (Menn et al., 2026b); local PBO (Menn et al., 2026a) | entered the field in 2026; tuning robot controllers | |
| Human-computer interaction and graphics (SIGGRAPH, UIST, CHI, IUI, Mensch und Computer, UMAP) | Koyama (University of Tokyo; National Institute of Advanced Industrial Science and Technology) | sequential line search (Koyama et al., 2017); Sequential Gallery (Koyama et al., 2020); constrained PBO (Iwai et al., 2025) | interactive optimization of visual design (Chapter 25), turning to constraints and language-model assistance |
| Oulasvirta (Aalto University) | (Colella et al., 2020; Chan et al., 2022) | human studies of human-in-the-loop optimization, mostly with ratings or performance feedback rather than comparisons | |
| Butz, Buschek, Mayer (LMU Munich, media informatics) | (Ou et al., 2022; Ou et al., 2023) | human-subject studies of PBO in creative tools | |
| Biomedical engineering (EMBC) | Fauvel, Chalk, Beyeler | (Fauvel and Chalk, 2021; Granley et al., 2023; Schoinas et al., 2025) | human-in-the-loop optimization of retinal prosthesis encoders |
Other recurring contributors include Mesbah and colleagues (KappaSharp, a 2026 preprint) (Shao et al., 2026), Wu and Gardner (a knowledge gradient for preference learning, a 2026 preprint) (Wu and Gardner, 2026), and Theiner, Hirt, Findeisen, and colleagues in control (ECC 2025 and ECC 2026) (Theiner et al., 2025; Theiner et al., 2026); we did not verify their institutions. Read across the table, the groups are moving from methods toward applications and toward language models as intermediaries, and the newcomers (Trimpe, Mesbah) come from control (inference). The groups that prove regret bounds do not run human studies, the groups that run human studies rarely change the acquisition function, and the groups that build the software sit closest to the decision-theoretic line (inference).
31.8.2 No survey, nine theses, no dedicated workshop #
There is no dedicated survey of PBO. General references on Bayesian optimization give preferences a page at most: Frazier's tutorial (Frazier, 2018) does not contain the words "preference", "pairwise", or "duel", Garnett's monograph (Garnett, 2023) mentions optimizing human preferences in about one paragraph, and the survey of Wang et al. (2023b) (ACM Computing Surveys 2023) has no section on PBO. The dueling bandit survey of Bengs et al. (2021) (JMLR 2021) covers preference probabilities modeled with Gaussian processes but cites neither González et al. 2017 nor Brochu et al. (Brochu et al., 2010). The closest reference work is the tutorial of Benavoli and Azzimonti (2026a) (Foundations and Trends in Machine Learning 2026), which covers nine Gaussian process models of preferences and choices but mentions PBO only briefly. The remaining teaching material consists of software tutorials and the book chapter Computational Design with Crowds by Koyama and Igarashi (Koyama and Igarashi, 2018), a preprint whose formal publication details we did not verify.
Doctoral theses carry much of the field's synthesis. We found at least nine doctoral theses from 2017 to 2025 that center on PBO or devote a substantial part to it (Table 31.7).
| Author | Title | Institution | Year |
|---|---|---|---|
| Yuki Koyama | Computational Design Driven by Visual Aesthetic Preference (Koyama, 2017) | University of Tokyo | 2017 |
| Ellen Novoseller | Online Learning from Human Feedback with Applications to Exoskeleton Gait Optimization (Novoseller, 2021) | Caltech | 2021 (record date) |
| Eero Siivola | Applications of human feedback in Gaussian processes (Siivola, 2021) | Aalto University | 2021 |
| Tristan Fauvel | Human-in-the-loop optimization of retinal prostheses encoders (Fauvel, 2021) | Sorbonne University | 2021 |
| Raul Astudillo | Exploiting Composite Functions in Bayesian Optimization (Astudillo Marban, 2022) | Cornell University | 2022 |
| Maegan Tucker | Enabling Robust and User-Customized Bipedal Locomotion on Lower-Body Assistive Devices via Hybrid System Theory and Preference-Based Learning (Tucker, 2023) | Caltech | 2023 |
| Petrus Mikkola | Humans as Information Sources in Bayesian Optimization (Mikkola, 2024) | Aalto University | 2024 |
| Mengjia Zhu | Global and preference-based optimization using surrogate-based methods (Zhu, 2024) | IMT School for Advanced Studies Lucca | 2024 |
| Wenjie Xu | Bayesian Optimization with Constraints, Structure and Human Feedback (Xu, 2025) | EPFL | 2025 |
The closest workshops covered preference learning as a whole. The ICML 2023 workshop The Many Facets of Preference-Based Learning covered dueling bandits, reinforcement learning from human feedback (RLHF, the method used to align language models with human preferences), social choice, and optimization, with two papers on Bayesian optimization from preferences (ICML, 2023). PBO work has also appeared at the ICML 2019 Workshop on Human in the Loop Learning (McCourt and Dewancker, 2019), at the NeurIPS 2022 Workshop on Gaussian Processes, Spatiotemporal Modeling, and Decision-making Systems (Takeno et al., 2022), and at ProbML 2026, a symposium held alongside ICML with archival proceedings (ProbML, 2026). No workshop dedicated to PBO appeared at NeurIPS or ICML from 2024 to 2026, as far as arXiv comments and three web searches show.
31.8.3 How much is published #
Publication volume is growing, from a small base. We ran four searches, two over arXiv abstracts (arXiv, 2026b; arXiv, 2026a) and two over Semantic Scholar (Semantic Scholar, 2026), each in a narrow form (the phrase) and a broad form (preference words together with Bayesian optimization). Figure 31.3 shows the yearly counts.
All four searches show growth by a factor of about 2.4 to 4.3 from 2023 to 2025, and the first nine months of 2026 nearly match all of 2025 (12 entries against 13, 38 against 39, 10 against 13, and 32 against 34). But even in 2025 and 2026 a year brings only 10 to 40 entries. The growth coincides with the attention that language-model alignment has brought to preference feedback (inference). The cross-citations we checked are few: the 2021 dueling bandit survey does not cite González et al. 2017, and the 2026 local PBO paper does not cite the work on high-dimensional lengthscale priors (Section 30.4) (inference).
Sources cited in Section 31.8 62
- Astudillo and Frazier (2020) Multi-attribute Bayesian optimization with interactive preference learning
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
- Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Ozaki et al. (2024) Multi-Objective Bayesian Optimization with Active Preference Learning
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- Benavoli and Azzimonti (2026a) A tutorial on learning from preferences and choices with Gaussian Processes
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Sinaga et al. (2026) Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization
- Kirschner and Krause (2021) Bias-Robust Bayesian Optimization via Dueling Bandits
- Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
- Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
- Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Bemporad and Piga (2021) Global optimization based on active preference learning with radial basis functions
- Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
- Zhu (2025) PWAS
- Tucker et al. (2020b) Preference-Based Learning for Exoskeleton Gait Optimization
- Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
- Tucker et al. (2022) POLAR: Preference Optimization and Learning Algorithms for Robotics
- Menn et al. (2026b) Preferential Bayesian Optimization with Crash Feedback
- Menn et al. (2026a) Local Preferential Bayesian Optimization
- Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Iwai et al. (2025) Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design
- Colella et al. (2020) Human Strategic Steering Improves Performance of Interactive Optimization
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Ou et al. (2023) The Impact of Expertise in the Loop for Exploring Machine Rationality
- Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization
- Granley et al. (2023) Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses
- Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Wu and Gardner (2026) Knowledge Gradient for Preference Learning
- Theiner et al. (2025) Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles
- Theiner et al. (2026) Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models
- Frazier (2018) A Tutorial on Bayesian Optimization
- Garnett (2023) Bayesian Optimization
- Wang et al. (2023b) Recent Advances in Bayesian Optimization
- Bengs et al. (2021) Preference-based Online Learning with Dueling Bandits: A Survey
- Brochu et al. (2010) A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning
- Koyama and Igarashi (2018) Computational Design with Crowds
- Koyama (2017) Computational Design Driven by Visual Aesthetic Preference
- Novoseller (2021) Online Learning from Human Feedback with Applications to Exoskeleton Gait Optimization
- Siivola (2021) Applications of human feedback in Gaussian processes
- Fauvel (2021) Human-in-the-loop optimization of retinal prostheses encoders
- Astudillo Marban (2022) Exploiting Composite Functions in Bayesian Optimization
- Tucker (2023) Enabling Robust and User-Customized Bipedal Locomotion on Lower-Body Assistive Devices via Hybrid System Theory and Preference-Based Learning
- Mikkola (2024) Humans as Information Sources in Bayesian Optimization
- Zhu (2024) Global and preference-based optimization using surrogate-based methods
- Xu (2025) Bayesian Optimization with Constraints, Structure and Human Feedback
- ICML (2023) The Many Facets of Preference-Based Learning
- McCourt and Dewancker (2019) Sampling Humans for Optimizing Preferences in Coloring Artwork
- Takeno et al. (2022) Preferential Bayesian Optimization with Hallucination Believer
- ProbML (2026) Symposium on Probabilistic Machine Learning website
- arXiv (2026b) Abstract search: preferential AND Bayesian AND (optimization OR optimisation)
- arXiv (2026a) Abstract search: preference terms AND "Bayesian optimization"
- Semantic Scholar (2026) Bulk search: "preferential bayesian optimization"
31.9 Settled, contested, missing #
Settled. The only maintained general-purpose PBO software is BoTorch and Ax, both from Meta; the other major Bayesian optimization and Gaussian process libraries have no preference support. None of the preference software we checked scales its default lengthscale prior with dimension. Research code has no release tags and outdated dependencies, and some of it has no license. Evaluation relies on simulated users and synthetic functions in 1 to 8 dimensions, with noise models, regret definitions, and numbers of repetitions chosen paper by paper. There is no dedicated survey of PBO and no workshop dedicated to it; the closest, at ICML 2023, covered preference-based learning as a whole. Publication volume has grown since 2023 from a small base.
Contested. Which acquisition function is best: the answer changes with dimension, noise, and metric, as qEUBO's better reported solution and worse cumulative regret against POP-BO show. How far results with simulated users carry over to people: the available evidence points to clear differences, but the studies are small.
Missing. A shared benchmark suite and leaderboard. A public data set of individual-level pairwise judgments. A reproduction study that reruns many methods under one protocol. A human-subject comparison with random assignment of acquisition functions. Tutorials for Ax's preference features and documentation of the default priors. A maintained PBO package for R or Julia (not searched systematically).
31.10 Exercises #
In Figure 31.2, each random query is a pair of two different points drawn uniformly from 41 candidates, and a run asks 16 such pairs. (a) What is the probability that a given candidate is never shown? (b) How many distinct candidates are shown on average? (c) On the running objective, three candidates have regret below 0.12: the optimum and its two neighbors. What is the probability that at least one of them is shown?
Solution
(a) A given candidate is left out of one pair with probability , so out of all 16 independent pairs with probability . (b) By linearity of expectation, distinct candidates. (c) Three given candidates are all left out of one pair with probability , and out of all 16 pairs with probability , so at least one is shown with probability about 0.915. Random pairs nearly exhaust a 41-point space; the best-queried metric rewards that coverage, and says nothing about whether the model learned where the optimum is.
The cumulative regret in Figure 31.2 adds, for every duel, the average regret of its two options, . Suppose that from some step on, the incumbent is the true optimum. What is the smallest possible per-duel regret of a rule that always includes the incumbent, and of a rule that never shows the same point twice? What does the metric reward?
Solution
A duel containing the optimum costs , at most half of the challenger's regret, and zero if the challenger is the optimum too. A rule that spreads its queries pays the full average regret of two new points, which on the running objective averages about 0.7 per point for a random candidate. Cumulative regret therefore rewards keeping the incumbent in every query, even when the challenger teaches the model little: it measures what the person was shown along the way, not how good the final recommendation is. That is why a method can be better on one metric and worse on another, as in the POP-BO comparison.
BOPE's simulated user picks the worse option 10% of the time, whatever the two options are. Under probit noise of size on each utility, the error probability for a utility gap is . (a) Find the for which a gap of is answered wrongly 10% of the time. (b) With that , how often is a gap of answered wrongly? (c) What does this say about comparing papers that use the two noise models?
Solution
(a) We need , so . (b) Then , and : almost never, against 10% under the flip model. (c) Under probit noise, errors concentrate on close calls; under a fixed flip rate, a clearly worse option is chosen as often as a marginally worse one. The second is harsher on methods that rely on a few decisive comparisons, so the same "10% noise" can mean very different difficulty, and results under the two models cannot be pooled (inference).
Further reading #
- The source of
PairwiseGP(Meta Platforms, Inc., 2026h) and BoTorch's PBO tutorial (Meta Platforms, Inc., 2026c) are the de facto definition of PBO in practice; read them before trusting a default. - Fauvel and Chalk (2021) is the one paper that evaluates many acquisition functions with formal multiple comparisons; its protocol is worth copying.
- Xu et al. (2024b) shows within one paper how the metric decides which method looks better.
- Schoinas et al. (2025) and Ou et al. (2022) are the clearest direct comparisons of simulated users and real people in preferential loops.
- Benavoli and Azzimonti (2026a) is the closest thing to a reference work on the models, though not on the optimization.
References
- (2022). PPBO. GitHub. software Cited in §31.1 §31.3
- (2026a). Abstract search: preference terms AND "Bayesian optimization". arXiv API. non-peer-reviewed Cited in §31.8
- (2026b). Abstract search: preferential AND Bayesian AND (optimization OR optimisation). arXiv API. non-peer-reviewed Cited in §31.8
- (2023a). qEUBO. GitHub. software Cited in §31.1 §31.3 §31.6
- (2023b). qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py). GitHub. software Cited in §31.4
- (2020). Multi-attribute Bayesian optimization with interactive preference learning. International Conference on Artificial Intelligence and Statistics. Cited in §31.8
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §31.4 §31.5 §31.6 §31.8
- (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Cited in §31.8
- (2022). Exploiting Composite Functions in Bayesian Optimization. Cornell University. thesis Cited in §31.8
- (2026). smac 2.4.1. PyPI. software Cited in §31.1
- (2023). GLIS. GitHub. software Cited in §31.1
- (2021). Global optimization based on active preference learning with radial basis functions. Machine Learning. Cited in §31.8
- (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. Cited in §31.8
- (2010). A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning. arXiv preprint. preprint Cited in §31.8
- (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. Cited in §31.7 §31.8
- (2020). Human Strategic Steering Improves Performance of Interactive Optimization. UMAP 2020. Cited in §31.7 §31.8
- (2023). preferentialBO. GitHub. software Cited in §31.1 §31.3 §31.6
- (2022). dragonfly-opt 0.1.7. PyPI. software Cited in §31.1
- (2026). preferential_batch_bayesian_optimization example. GitHub. software Cited in §31.1
- (2022). ax-platform 0.2.6. PyPI. software Cited in §31.1
- (2021). Human-in-the-loop optimization of retinal prostheses encoders. Sorbonne Université. thesis Cited in §31.8
- (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. preprint Cited in §31.4 §31.8
- (2017). candy-power-ranking data. GitHub. non-peer-reviewed Cited in §31.5 §31.6
- (2018). A Tutorial on Bayesian Optimization. arXiv. preprint Cited in §31.8
- (2023). Bayesian Optimization. Cambridge University Press. Cited in §31.8
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §31.4
- (2026). gpflow 2.11.1. PyPI. software Cited in §31.1
- (2026). gpytorch 1.15.2. PyPI. software Cited in §31.1
- (2023). Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. NeurIPS 2023. Cited in §31.7 §31.8
- (2024). HEBO 0.3.6. PyPI. software Cited in §31.1
- (2023). The Many Facets of Preference-Based Learning. ICML 2023 workshop page. non-peer-reviewed Cited in §31.8
- (2026). crashpbo. GitHub. software Cited in §31.1
- (2025). Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. Cited in §31.8
- (2026). SUSHI Preference Data Sets. kamishima.net. non-peer-reviewed Cited in §31.5
- (2025). BOHF_code_submission. GitHub. software Cited in §31.1
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §31.4 §31.8
- (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. Cited in §31.8
- (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Cited in §31.4 §31.8
- (2017). Computational Design Driven by Visual Aesthetic Preference. The University of Tokyo. doi:10.15083/00076184. thesis Cited in §31.8
- (2025a). preference-regressor.hpp. GitHub. software Cited in §31.2
- (2025b). sequential-line-search. GitHub. software Cited in §31.1
- (2018). Computational Design with Crowds. Computational Interaction. Cited in §31.8
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Cited in §31.2 §31.8
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Cited in §31.2 §31.8
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §31.4 §31.8
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §31.4 §31.6 §31.8
- (2019). Sampling Humans for Optimizing Preferences in Coloring Artwork. ICML 2019 Workshop on Human in the Loop Learning. workshop paper Cited in §31.8
- (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Cited in §31.4 §31.8
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Cited in §31.8
- (2026). AEPsych. GitHub. software Cited in §31.1 §31.3
- (2026a). ax-platform release history. PyPI. software Cited in §31.1
- (2026b). ax/generation_strategy/transition_criterion.py. GitHub. software Cited in §31.1
- (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. software Cited in §31.1
- (2026e). BoTorch CHANGELOG. GitHub. software Cited in §31.1
- (2026f). BoTorch LICENSE. GitHub. software Cited in §31.1
- (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Cited in §31.2
- (2026i). botorch release history. PyPI. software Cited in §31.1
- (2026j). botorch/acquisition/preference.py. GitHub. software Cited in §31.1
- (2026k). botorch/models/utils/gpytorch_modules.py. GitHub. software Cited in §31.2
- (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Cited in §31.1
- (2026m). tutorials directory. GitHub. software Cited in §31.1
- (2023). qEUBO. GitHub. software Cited in §31.1 §31.3 §31.6
- (2026). lilo. GitHub. software Cited in §31.1
- (2024). Humans as Information Sources in Bayesian Optimization. Aalto University. thesis Cited in §31.8
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §31.4 §31.7 §31.8
- (2021). Online Learning from Human Feedback with Applications to Exoskeleton Gait Optimization. California Institute of Technology. doi:10.7907/gvtx-1586. thesis Cited in §31.8
- (2025). Ax: A Platform for Adaptive Experimentation. International Conference on Automated Machine Learning. Cited in §31.1
- (2026a). optuna 5.0.0. PyPI. software Cited in §31.1
- (2026b). optuna-dashboard 0.21.0. PyPI. software Cited in §31.1
- (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. software Cited in §31.2
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §31.7 §31.8
- (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. Cited in §31.7 §31.8
- (2026). PLMBO (Preference Learning Multi-Objective Bayesian Optimization). OptunaHub. software Cited in §31.1
- (2024). Multi-Objective Bayesian Optimization with Active Preference Learning. Proceedings of the AAAI Conference on Artificial Intelligence. Cited in §31.1 §31.8
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Cited in §31.4 §31.8
- (2024). POP-BO. GitHub. software Cited in §31.1
- (2026). Symposium on Probabilistic Machine Learning website. probml.cc. non-peer-reviewed Cited in §31.8
- (2026). DT-PBO-preprint. GitHub. software Cited in §31.1
- (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Cited in §31.7 §31.8
- (2026). trieste 4.6.0. PyPI. software Cited in §31.1
- (2026). Bulk search: "preferential bayesian optimization". Semantic Scholar API. non-peer-reviewed Cited in §31.8
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §31.4 §31.8
- (2023). GPyOpt (archived). GitHub. software Cited in §31.1
- (2021). Applications of human feedback in Gaussian processes. Aalto University. thesis Cited in §31.8
- (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Cited in §31.4 §31.6 §31.8
- (2026). Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization. Symposium on Probabilistic Machine Learning (ProbML 2026), Proceedings Track. Cited in §31.8
- (2026). Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis. arXiv. preprint Cited in §31.7
- (2022). Preferential Bayesian Optimization with Hallucination Believer. NeurIPS 2022 Workshop on Gaussian Processes, Spatiotemporal Modeling, and Decision-making Systems. workshop paper Cited in §31.8
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §31.4 §31.8
- (2025). Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles. ECC 2025. Cited in §31.8
- (2026). Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models. European Control Conference. Cited in §31.8
- (2023). Enabling Robust and User-Customized Bipedal Locomotion on Lower-Body Assistive Devices via Hybrid System Theory and Preference-Based Learning. California Institute of Technology. doi:10.7907/j9hk-xa17. thesis Cited in §31.8
- (2024). POLAR. GitHub. software Cited in §31.1
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Cited in §31.8
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Cited in §31.8
- (2022). POLAR: Preference Optimization and Learning Algorithms for Robotics. arXiv. preprint Cited in §31.8
- (2023b). Recent Advances in Bayesian Optimization. ACM Computing Surveys. Cited in §31.8
- (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Cited in §31.8
- (2025). Bayesian Optimization with Constraints, Structure and Human Feedback. École Polytechnique Fédérale de Lausanne (EPFL). doi:10.5075/epfl-thesis-11166. thesis Cited in §31.8
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §31.4 §31.8
- (2025). PABBO code repository: evaluation config evaluate.yaml. GitHub. software Cited in §31.2
- (2026). PABBO. GitHub. software Cited in §31.1
- (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §31.2 §31.4 §31.6 §31.8
- (2024). Global and preference-based optimization using surrogate-based methods. IMT School for Advanced Studies Lucca. doi:10.13118/imtlucca/e-theses/415. thesis Cited in §31.8
- (2025). PWAS. GitHub. software Cited in §31.1 §31.8
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Cited in §31.8