Bayesian Optimization
Part VI: The Research Frontier
中文

Observation Models, Surrogates, and Inference

Every preferential Bayesian optimization (PBO) system rests on three modeling choices, which Part IV introduced one at a time. The observation model says how a comparison, a ranking, or a choice arises from the latent utility: in Chapter 16, a probit or logistic link applied to a utility difference. The surrogate is the prior over the utility: in Chapter 18, a Gaussian process. Posterior inference combines the two and hands the result to the acquisition function: in Chapter 17, usually the Laplace approximation.

From 2017 to September 2026 the default combination barely changed. A Gaussian process prior, a probit link, and the Laplace approximation, which is the model of Chu and Ghahramani (2005), remain what BoTorch's PairwiseGP implements (Meta Platforms, Inc., 2026h). What changed is the scrutiny each layer received, and this chapter follows the three layers in order, ending with what the default implementation decides for you.

Sources cited in the introduction 2
  1. Chu and Ghahramani (2005) Preference learning with Gaussian processes
  2. Meta Platforms, Inc. (2026h) BoTorch PairwiseGP source code pairwise_gp.py

27.1 The baseline models #

Chu and Ghahramani (2005) placed a Gaussian process prior on the latent utility ff and used the likelihood

P(x≻x′)=Φ ⁣(f(x)−f(x′)2 σ),\Prob(\vx \succ \vx') = \Phi\!\left(\frac{f(\vx) - f(\vx')}{\sqrt{2}\,\sigma}\right),
(27.1)

which is Thurstone's comparative judgment with Gaussian noise of scale σ\sigma on each option (Section 16.3). They fitted the posterior with the Laplace approximation and chose hyperparameters by the approximate evidence. The pairwise model of Brochu et al. (2007) is the same Thurstone model, as Koyama et al. (2020) point out. Houlsby et al. (2011) took another route to the same place: a preference kernel that turns pairwise preference learning into Gaussian process classification over pairs.

The paper that named the field used neither the probit link nor, as far as its text says, the Laplace approximation. González et al. (2017) modeled the probability that x\vx wins a duel against x′\vx' with a Gaussian process classifier on the dueling space X×X\X \times \X, with a Bernoulli likelihood and the latent joint reward squashed by a logistic function, assuming the latent function has the form f([x,x′])=g(x′)−g(x)f([\vx, \vx']) = g(\vx') - g(\vx); their text says only that the posterior is intractable and needs approximations, without naming one. They listed the doubled input dimension as a limitation, and Nguyen et al. (2021) later criticized the model for working in twice the dimension of the objective, for not modeling the objective directly, and for not modeling ties. Because the skew Gaussian process theorem of Section 27.4 concerns the probit likelihood, it does not apply directly to this original model (inference).

Since then the two links have been used side by side. BoTorch offers both: PairwiseProbitLikelihood with P(v≻u)=Φ((f(v)−f(u))/2)\Prob(v \succ u) = \Phi((f(v) - f(u))/\sqrt{2}) and PairwiseLogitLikelihood with sigmoid⁡(f(v)−f(u))\operatorname{sigmoid}(f(v) - f(u)) (Meta Platforms, Inc., 2026g). The logistic link appears in the noise analysis of qEUBO (Astudillo et al., 2023), in POP-BO (Xu et al., 2024b), and in MR-LPF (Kayal et al., 2025). For two options, a softmax over utilities is exactly the logistic Bradley-Terry model. Both links are random utility models (Section 16.5): the probit one comes from Gaussian noise on each option, and the logistic one from Gumbel noise, a correspondence Kayal et al. use when they discuss lower bounds.

Key idea The link was inherited, not chosen

We found no PBO paper that compares the probit and logistic links on human data, and no study of robustness when the link is wrong. The only data-driven alternative comes from Shvartsman et al. (2024): a drift-diffusion model of the decision process induces its own choice link, but the gains they measured appeared after adding response-time data and cannot be credited to the link itself (inference). The choice of link is a matter of lineage and convenience, not of evidence (inference). Until a comparison exists, a practitioner can treat the link as a checkable choice: fit both to the answers collected and compare how well each predicts held-out comparisons (inference).

Sources cited in Section 27.1 11
  1. Chu and Ghahramani (2005) Preference learning with Gaussian processes
  2. Brochu et al. (2007) Active Preference Learning with Discrete Choice Data
  3. Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
  4. Houlsby et al. (2011) Bayesian Active Learning for Classification and Preference Learning
  5. González et al. (2017) Preferential Bayesian Optimization
  6. Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
  7. Meta Platforms, Inc. (2026g) BoTorch pairwise likelihood source code likelihoods/pairwise.py
  8. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  9. Xu et al. (2024b) Principled Preferential Bayesian Optimization
  10. Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
  11. Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences

27.2 Extending the observation model #

A person's answer to "which is better?" can carry more, or different, information than one bit. Since 2017, researchers have extended the likelihood in eight directions. Table 27.1 summarizes them; the paragraphs after it give the details that matter for judging each one.

Table 27.1 Extensions of the observation model since 2017, and the human evidence for each.
Extension Representative work Likelihood Human data
Choice from a set, batch winner, ranking (Koyama et al., 2017; Koyama et al., 2020; Siivola et al., 2021; Nguyen et al., 2021; Benavoli et al., 2023) Bradley-Terry for mm options; batch-winner likelihood; multinomial logit and top-kk rankings; choice functions with several utilities crowdsourcing and a 6-person pilot; benchmarks built from rating data; simulations only for Benavoli et al.
Ties and indifference (Bıyık et al., 2019) (linear reward); (Nguyen et al., 2021; Benavoli and Azzimonti, 2026a; Erarslan et al., 2025) just-noticeable-difference threshold; multinomial logit with ties; indistinguishability ∣u(x)−u(x′)∣/δ≤1\lvert u(\vx) - u(\vx')\rvert / \delta \le 1 only Bıyık et al.'s user study, with a linear model
Ordinal labels and confidence ROIAL (Li et al., 2021); (Dao et al., 2025; Wu et al., 2025a; Peng et al., 2025; Zhang et al., 2026b) ordered thresholds; 5-level Likert strength; Likert confidence with cut points and lapse rates 3 participants (ROIAL); offline fits (Wu et al.); a user study (Peng et al.); N=13N = 13 (Zhang et al.)
Response times (Shvartsman et al., 2024); (Li et al., 2024a) (linear) differentiable approximation to the drift-diffusion likelihood three data sets, offline
Failure, validity, crashes (Benavoli et al., 2021c); C-GLISp (Zhu et al., 2022); CrashPBO (Menn et al., 2026b) valid/invalid label joint with preference; feasibility surrogate; crash as a second outcome controller calibration; three robot platforms
Intransitivity (Chau et al., 2022) Gaussian process on a skew-symmetric preference function animal contests, NFL games, citations; no design preferences
Heteroscedastic noise (Sinaga et al., 2026) input-dependent noise from a kernel density over anchors Sushi and Candy benchmarks, no live users
Several users, several utilities (Houlsby et al., 2012; Simpson and Gurevych, 2020; Benavoli et al., 2023; Astudillo et al., 2025; Dubey et al., 2026) collaborative low-rank GP; matrix factorization with GPs; several latent utilities; Dirichlet process mixture real multi-user data (Houlsby, Simpson); simulations otherwise
Utility drifting over time none none none

Choices from sets, and rankings. The sequential line search of Koyama et al. (2017) records the person's pick along a slider as a preference for the chosen point over both the current best and the point of highest expected improvement, using the Bradley-Terry-Luce model, the extension of Bradley-Terry to a choice among mm options; the Sequential Gallery models pairwise data with the Thurstone model and choices among mm options with the Bradley-Terry-Luce form (Koyama et al., 2020). Siivola et al. (2021) give a likelihood for arbitrary parallel feedback on two or more points, focus on the batch winner, and argue that full rankings are laborious for people. Nguyen et al. (2021) build a surrogate inspired by the multinomial logit and its ranking extension, interpretable as Gaussian process regression with i.i.d. Gumbel noise, which handles top-kk rankings and ties. Benavoli et al. (2023) handle set-valued choices (picking, say, three of five options shown, when none of the three is clearly better than the others) with several latent utilities; their abstract reports only simulations.

Ties and indifference. Bıyık et al. (2019) added an "About Equal" answer with a minimum perceivable difference δ≥0\delta \ge 0: P(About Equal)=(e2δ−1) P(choose a) P(choose b)\Prob(\text{About Equal}) = (e^{2\delta} - 1)\,\Prob(\text{choose }a)\,\Prob(\text{choose }b), which reduces to the softmax model at δ=0\delta = 0. Their model is a linear reward, not a Gaussian process, and comes with simulations and a user study. C-GLISp encodes better, worse, or similar as constraints on a radial basis function surrogate (Zhu et al., 2022). The tutorial of Benavoli and Azzimonti lists nine likelihoods, among them a just-noticeable-difference model that adds statements of indistinguishability, ∣u(x)−u(x′)∣/δ≤1\lvert u(\vx) - u(\vx')\rvert/\delta \le 1, going back to Luce's 1956 notion of a discrimination threshold. Erarslan et al. (2025) represent indifference with a just-noticeable-difference threshold in an extended Thurstone model and report clear gains when 10% to 20% of comparisons are indifferent. The field has converged on two treatments, a threshold parameter in a Thurstone or logit model, or ties in a multinomial logit (inference). We found no Gaussian process PBO paper with an "abstain" or "skip" outcome distinct from a tie, the answer a person gives when they cannot compare the options at all.

Ordinal labels, strength, and confidence. ROIAL (Li et al., 2021) combines pairwise preferences with ordinal labels (very bad, bad, neutral, good), arguing that an rr-level ordinal query yields at most log⁡2r\log_2 r bits against one bit for a preference; its participants were 3 non-disabled people tuning 4 exoskeleton gait parameters. Wu et al. (2025a) add to the probit preference likelihood a Likert confidence likelihood with learnable cut points and a lapse rate; on human comparisons of robot gaits, they report that models trained with the confidence ratings "consistently achieve lower Brier scores and higher F1 scores", that is, smaller squared errors of the predicted probabilities and more accurate predicted choices.

Response times. How long a person takes to answer says something about how different the options felt. Shvartsman et al. (2024) approximated the drift-diffusion likelihood with a moment-matched skewed three-parameter distribution so that response times can enter a variational Gaussian process, and describe this as the first Gaussian process model with a joint likelihood of drift-diffusion response times and choices. On one participant's judgments of 1,225 pairs of quadruped robot gaits and on a visual psychophysics data set, the joint drift-diffusion models predicted held-out choices significantly better than the choice-only model, especially with small training sets, while a simpler stacking approach, which gives the choice model the logarithm of the response time as an extra input, did not improve on choice alone; on a recommender-system data set every joint model improved, stacking most. Li et al. (2024a) used the EZ-diffusion model for fixed-budget best-arm identification with linear utilities and proved that, for strong preferences, response times add information about preference strength.

Failure, validity, and crashes. Some experiments produce no output at all. Validity labels and feasibility surrogates attach a second label to each sample, and CrashPBO (Menn et al., 2026b) treats a crash report as a second kind of outcome, reduces crashes by 63% on synthetic benchmarks, and was validated on three robot platforms.

Intransitivity. Chau et al. (2022) placed a Gaussian process on a skew-symmetric preference function g(x,x′)g(\vx, \vx') with a kernel they call the generalized preferential kernel, which can represent cycles. On real data the new model's accuracy against the Chu-Ghahramani model was 0.78 against 0.51 for chameleon contests, 0.83 against 0.80 for flat lizards, 0.59 against 0.51 for NFL games from 2000 to 2018 (where linear pairwise logistic regression did best, at 0.65), and 0.74 against 0.66 for an arXiv citation graph. The authors conclude that their findings support the conjecture that violations of rankability, the assumption that one utility orders all items, are ubiquitous in real preference data. These data sets are ones where intransitivity is expected (contests and sports), the linear model won on the only human competition data, and none is a single user's design preference, so carrying the conclusion over to typical PBO settings lacks support (inference).

Heteroscedastic noise. Some comparisons are harder than others. Sinaga et al. (2026) ask the user for reliable "anchor" points and turn them into an input-dependent noise map with kernel density estimation, used in the probit likelihood Φ((f(x)−f(x′))/σ2(x)+σ2(x′))\Phi\big((f(\vx) - f(\vx'))/\sqrt{\sigma^2(\vx) + \sigma^2(\vx')}\big) or a logistic likelihood with input-dependent temperature; their consistency analysis holds only under an idealized i.i.d. anchor model, and the evaluation uses the Sushi and Candy data without live users. Fauvel and Chalk (2021) argue that acquisition functions for binary and preferential Bayesian optimization must separate epistemic from aleatoric uncertainty, since only the former should drive exploration.

Several users, and drift. Models for many users share low-dimensional structure across them, and crowdGPPL (Simpson and Gurevych, 2020) scales a combination of matrix factorization and Gaussian processes to thousands of users and items; Section 20.5 describes these models. We found no PBO or Gaussian process preference paper that models a utility changing over time; the finite-arm theory that does exist is in Section 29.10.

What the human evidence supports. For ties there is one user study with a linear model; for ordinal labels, three participants; for confidence, one offline robot-gait data set and two small user studies; for response times, three data sets, few participants, and offline fits; for crashes, three robot platforms; for heteroscedastic noise, two benchmarks converted from rating data; for intransitivity, no design preferences at all. Most evaluations compare model fits offline rather than closing the loop with people (inference). Of the signals that cost the person nothing extra, response times and confidence are the two that have improved models on real human data, which makes them the best-supported upgrades; recording both in every session costs nothing and lets them be tested in closed-loop PBO later (inference).

Sources cited in Section 27.2 25
  1. Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
  2. Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
  3. Siivola et al. (2021) Preferential Batch Bayesian Optimization
  4. Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
  5. Benavoli et al. (2023) Learning Choice Functions with Gaussian Processes
  6. Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
  7. Benavoli and Azzimonti (2026a) A tutorial on learning from preferences and choices with Gaussian Processes
  8. Erarslan et al. (2025) Consecutive Preferential Bayesian Optimization
  9. Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
  10. Dao et al. (2025) Experience in Engineering Complex Systems: Active Preference Learning With Multiple Outcomes and Certainty Levels
  11. Wu et al. (2025a) Mixed Likelihood Variational Gaussian Processes
  12. Peng et al. (2025) Towards Uncertainty Unification: A Case Study for Preference Learning
  13. Zhang et al. (2026b) Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
  14. Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences
  15. Li et al. (2024a) Enhancing Preference-based Linear Bandits via Human Response Time
  16. Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
  17. Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
  18. Menn et al. (2026b) Preferential Bayesian Optimization with Crash Feedback
  19. Chau et al. (2022) Learning Inconsistent Preferences with Gaussian Processes
  20. Sinaga et al. (2026) Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization
  21. Houlsby et al. (2012) Collaborative Gaussian Processes for Preference Learning
  22. Simpson and Gurevych (2020) Scalable Bayesian preference learning for crowds
  23. Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
  24. Dubey et al. (2026) Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization
  25. Fauvel and Chalk (2021) Efficient Exploration in Binary and Preferential Bayesian Optimization

27.3 Surrogates beyond the Gaussian process #

The Gaussian process utility remains the workhorse, but four families of alternatives appeared, each for a specific reason.

Gaussian process variants. Bıyık et al. (2020) fixed f(Ψˉ)=0f(\bar{\Psi}) = 0 at a reference point through the kernel, removing the additive non-identifiability of Section 18.4 in the prior. The skew Gaussian process line was built by Benavoli, Azzimonti, and Piga: a 2020 classification paper introduced the skew Gaussian process and proved it conjugate to the probit likelihood (Benavoli et al., 2020), a 2021 paper extended conjugacy to normal and affine probit likelihoods and their products (Benavoli et al., 2021a), and Benavoli and Azzimonti (2024) showed that a Gaussian process with linear inequality constraints is a skew Gaussian process, using it for monotone preference learning.

Radial basis functions, neural networks, and trees. GLISp (Bemporad and Piga, 2021) fits a radial basis function surrogate by linear or quadratic programming so that it satisfies the observed preferences as far as possible, and reports that, with the same number of comparisons, it usually gets closer to the optimum than PBO at lower computational cost; GLISp-r (Previtali et al., 2023) adds a proof of global convergence. These surrogates are not probabilistic, so the acquisition functions of Chapter 19 do not apply to them. Neural dueling bandits (Verma et al., 2025) estimate rewards from preference feedback with a neural network and prove sublinear regret, on synthetic data, and Wang et al. (2025a) replace the Gaussian process utility with an ensemble of monotone neural networks in preference exploration. DT-PBO (Leenders et al., 2025) builds shallow decision trees directly from comparisons, with a Laplace approximation in each leaf, and its authors report convergence competitive with Gaussian process PBO on 8 benchmark functions, particularly on rugged ones.

Amortized surrogates. PABBO (Zhang et al., 2025a) starts from the observation that the non-conjugate likelihood makes every PBO step expensive. It meta-learns surrogate and acquisition together in a transformer neural process, trained with reinforcement learning plus an auxiliary preference-prediction loss on synthetic tasks such as Gaussian process samples. Its abstract claims it is "several orders of magnitude faster than the usual Gaussian process-based strategies"; the paper reports a 12-fold average speedup over the fastest Gaussian process strategy on five problems and about 10-fold over noisy batch expected improvement on Gaussian process tasks. Its default evaluation configuration uses noise-free comparisons (Zhang, 2025). The authors list as limitations that the query set grows with dimension, each dimension needs its own model, how to build the pretraining data is open, and the person making comparisons is not modeled.

Priors across users, and generative latent spaces. Cross-user structure has been implemented through collaborative models such as crowdGPPL; through transfer at the acquisition level, as in Meta-PO (Li et al., 2025a), which stores earlier users' sessions as separate Gaussian processes and transfers them with a rank-weighted acquisition function (Section 32.5 reports its user study); and through guidance models trained on population data, such as the deep encoder of Granley et al. (2023). Several systems also optimize in the latent space of a generative model, such as FontCraft (Tatsukawa et al., 2025) in a font-style space and GimmBO (Liu et al., 2026b) over the merging weights of 20 to 30 diffusion adapters; none of these papers models how the geometry of the latent space relates to the utility's lengthscale (inference).

Overall, the debate about surrogates since 2020 has been mainly about the quality of inference in the same model, Gaussian approximation against the exact skew Gaussian process, not about the model class. Changes of model class were driven by structure (monotonicity, intransitivity, several utilities) or speed (PABBO, GLISp), not by evidence that a Gaussian process utility prior is wrong for human preferences (inference).

Sources cited in Section 27.3 15
  1. Bıyık et al. (2020) Active Preference-Based Gaussian Process Regression for Reward Learning
  2. Benavoli et al. (2020) Skew Gaussian processes for classification
  3. Benavoli et al. (2021a) A unified framework for closed-form nonparametric regression, classification, preference and mixed problems with Skew Gaussian Processes
  4. Benavoli and Azzimonti (2024) Linearly Constrained Gaussian Processes are SkewGPs: application to Monotonic Preference Learning and Desirability
  5. Bemporad and Piga (2021) Global optimization based on active preference learning with radial basis functions
  6. Previtali et al. (2023) GLISp-r: a preference-based optimization algorithm with convergence guarantees
  7. Verma et al. (2025) Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
  8. Wang et al. (2025a) Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble
  9. Leenders et al. (2025) DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization
  10. Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
  11. Zhang (2025) PABBO code repository: evaluation config evaluate.yaml
  12. Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
  13. Granley et al. (2023) Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses
  14. Tatsukawa et al. (2025) FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization
  15. Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization

27.4 Posterior inference #

Chapter 17 introduced the approximations, and Section 17.7 reported the main evidence on how well they do; this section adds the details behind that evidence and the choices the libraries make. Table 27.2 lists where each method is used.

Table 27.2 Inference methods for preference models: where they are used and the problems recorded.
Method Used by Recorded problems
Laplace approximation Chu and Ghahramani 2005; BoTorch PairwiseGP; Bıyık et al. 2020; ROIAL 2021; projective PBO 2020 the mode can be far from the mean, worse at low noise; rank-deficient likelihood Hessian with EUBO (preprint)
Expectation propagation Houlsby et al. 2012 (with variational Bayes); Siivola et al. 2021; Fauvel and Chalk 2021; optuna-dashboard accurate means and credible intervals; duel probabilities near 0 or 1 distorted, relative order not preserved
Variational inference Simpson and Gurevych 2020; Nguyen et al. 2021; Benavoli et al. 2023; Shvartsman et al. 2024; Wu et al. 2025; qEUBO experiments no systematic error assessment for preference likelihoods
Exact skew GP sampling Benavoli et al. 2021 (LinESS); Takeno et al. 2023 (Gibbs) every prediction needs truncated multivariate normal sampling; the evidence needs high-dimensional normal CDFs
Hallucination believer Takeno et al. 2023 a single truncated-posterior sample; stuck in local optima under logistic noise (reported by POP-BO)
Amortization PABBO 2025 fixed dimension; depends on pretraining data

The Laplace approximation is the library default, mostly for speed. Bıyık et al. (2020) wrote of expectation propagation that "while it is more accurate than Laplace approximation, it is slower in practice", and chose Laplace; projective PBO (Mikkola et al., 2020) uses it "for the sake of simplicity". Koyama et al. use maximum a posteriori estimates with a tight log-normal prior on each hyperparameter, written LN(μ,0.01)\mathrm{LN}(\mu, 0.01) in the paper, with μ=0.5\mu = 0.5 for every lengthscale and μ=0.2\mu = 0.2 for the amplitude, which nearly fixes the lengthscale (Koyama et al., 2020). Expectation propagation is the inference method in Siivola et al.'s experiments on real data, and the preferential Gaussian process sampler of optuna-dashboard uses it to fit its hyperparameters (Optuna developers, 2026c).

The exact posterior is a skew Gaussian process. Benavoli et al. (2021c) proved this in their PBO paper (Theorem 29.1) and sampled the posterior with LinESS, rejection-free elliptical slice sampling for linearly truncated Gaussians, at a cost they give as O(n3)O(n^3) time and O(n2)O(n^2) memory, the same bottleneck (a Cholesky factorization) as an ordinary Gaussian process; they report consistently better convergence and computation time than Laplace-based PBO.

How wrong the Gaussian approximations are. Takeno et al. (2023) took Gibbs sampling as ground truth (10,000 samples after 1,000 burn-in, thinned by 10), with an RBF kernel, noise variance 10−410^{-4}, and uniformly random duels. Beyond the findings of Section 17.7, that the Laplace mode can be very far from the mean at small noise and that expectation propagation is accurate for means and credible intervals but distorts duel probabilities, they report that on the Ackley function expectation propagation underestimates probabilities near 0 or 1 and does not preserve the true ordering, which affects acquisition functions that depend on the joint distribution of two points, and that Gibbs sampling mixed faster than LinESS. They consider unrealistic the setting in which Benavoli et al. compared expectation propagation with the skew Gaussian process, where all inputs in one interval lose and all inputs in another win; an earlier study of binary Gaussian process classification had found expectation propagation accurate (Kuss and Rasmussen, 2005). And they reject Laplace for the posterior yet use the Laplace evidence for hyperparameters, so the accuracy needed for hyperparameters and for acquisition is being treated separately.

Takeno et al.'s strongest evidence comes from noise variance 10−410^{-4}, where the posterior is nearly a truncated Gaussian; with noisier comparisons, closer to people, the skew should weaken. As of September 2026, no paper measures how large the errors of Laplace and expectation propagation are at realistic human noise levels, or replicates Takeno et al.'s calibration on real human comparisons. The experiment is cheap to describe: take comparisons from real people, run exact sampling as ground truth, and measure how far each approximation's predicted duel probabilities fall from it.

Sources cited in Section 27.4 7
  1. Bıyık et al. (2020) Active Preference-Based Gaussian Process Regression for Reward Learning
  2. Mikkola et al. (2020) Projective Preferential Bayesian Optimization
  3. Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
  4. Optuna developers (2026c) optuna-dashboard PreferentialGPSampler source code gp.py
  5. Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
  6. Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
  7. Kuss and Rasmussen (2005) Assessing Approximate Inference for Binary Gaussian Process Classification

27.5 Inside the default implementation #

Most people who run PBO run BoTorch's PairwiseGP, so its defaults are, in effect, the field's defaults. Table 18.2 listed the choices inside version 0.18.1 (Meta Platforms, Inc., 2026h; Meta Platforms, Inc., 2026g). Three of them shape what the model can learn. The probit likelihood is Φ((f(v)−f(u))/2)\Phi((f(v) - f(u))/\sqrt{2}), which fixes the noise at 1, so the output scale (BoTorch's name for the kernel's amplitude σf2\sigma_f^2) acts as the inverse of the noise level: it and the lengthscale must be learned jointly from comparisons alone, with no separate noise parameter, under a smoothed box prior on [0.01,100][0.01, 100] and a constraint of [0.005,200][0.005, 200] that a source comment calls a rule of thumb to keep estimates away from scales that saturate Φ\Phi. The probit argument is clipped to [−3,3][-3, 3], which caps the likelihood of any single comparison at about 0.9987 during fitting, an implicit robustness to mislabeled answers (Section 18.6). And the lengthscale prior is Gamma(2.4, 2.7), initialized at its mode, in every dimension.

The changelog adds the history (Meta Platforms, Inc., 2026e). Version 0.9.0 (August 2023) added a pairwise Bayesian active learning by disagreement acquisition function; 0.10.0 (February 2024) added qEUBO; 0.12.0 (September 2024) switched most models to lengthscale priors that scale with dimension but explicitly excluded PairwiseGP; 0.18.0 (June 2026) fixed PairwiseGP's state handling in evaluation and cross-validation; and in 0.18.1 (June 2026) PairwiseGP still uses Gamma(2.4, 2.7). The dimension-scaled priors came from a study of scalar Bayesian optimization by Hvarfner et al. (2024), the subject of Chapter 30. The preferential sampler in optuna-dashboard, based on Takeno et al., uses a Matérn 3/2 kernel with one lengthscale per dimension and a Gamma(5, 10) prior (mode 0.4), also independent of dimension (Optuna developers, 2026c).

What the fixed prior does in high dimension can be computed (inference). The mode of Gamma(2.4, 2.7) is (2.4−1)/2.7≈0.52(2.4 - 1)/2.7 \approx 0.52 in every dimension, while the dimension-scaled prior's mode on the unit cube grows from about 0.65 at 10 dimensions to about 2.05 at 100. Since random points in [0,1]d[0, 1]^d lie about d/6\sqrt{d/6} apart (Section 30.1.2 derives this), the default preference model in high dimension works where kernel values between typical points are near zero, and every comparison is nearly uninformative about every other point. This has not been verified directly on PairwiseGP.

Sources cited in Section 27.5 5
  1. Meta Platforms, Inc. (2026h) BoTorch PairwiseGP source code pairwise_gp.py
  2. Meta Platforms, Inc. (2026g) BoTorch pairwise likelihood source code likelihoods/pairwise.py
  3. Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
  4. Hvarfner et al. (2024) Vanilla Bayesian Optimization Performs Great in High Dimensions
  5. Optuna developers (2026c) optuna-dashboard PreferentialGPSampler source code gp.py

27.6 Comparison graphs and conditioning #

The numerical trouble reported in 2026 has a linear-algebra core that Section 18.5 derived: the Hessian W\mW of the negative log-likelihood is the weighted graph Laplacian of the comparison graph, whose nodes are compared inputs and whose edges are answered pairs, and it has one zero eigenvalue for every connected component. The Laplace posterior precision K−1+W\mK^{-1} + \mW is always full rank, because the prior contributes K−1\mK^{-1}, so the practical problem is not singularity but poor conditioning when the prior is weak along the shift directions, that is, when lengthscales are long or the output scale is large. The same fact has a statistical face, which Figure 27.1 shows: a shift direction that no comparison constrains is a relative utility that no answer has measured, and only the prior can say anything about it.

compared inputs on [0, 1]6 comparisons · 12 inputs · 6 components01234eigenvalues of the likelihood Hessian Wglobal shift (harmless)5 unmeasured offsets between groups
compared inputs on [0, 1]6 comparisons · 12 inputs · 6 components01234eigenvalues of the likelihood Hessian Wglobal shift (harmless)5 unmeasured offsets between groups
Figure 27.1 What the comparison graph lets the answers determine. Nodes are compared inputs, edges are answered duels, colors are connected components. Isolated pairs mimics queries that share no input with earlier ones, as EUBO tends to choose; Chain, Star, and Random pairs spend the same number of comparisons. Below, the eigenvalues of the likelihood Hessian W (unit curvature per comparison). Every zero is a direction no answer constrains: the first is the global shift, which never matters, and each further one (red) is an offset between two groups of inputs that were never compared with each other, which only the prior can fill in. The designs are illustrative.

Shao et al. (2026), a 2026 preprint, report that in the pipeline of PairwiseGP, the Laplace approximation, and EUBO, the new pairs EUBO selects share no candidate with earlier queries, so each pair forms an isolated component of the comparison graph and the likelihood Hessian becomes rank deficient. They call the deficiency structural, one that "cannot be resolved by changing the surrogate modeling approach", and note that existing remedies either force comparisons to stay connected, wasting query budget, or apply uniform regularization that perturbs well-constrained directions. Their correction, whose size Section 18.5 reports, adds to the Hessian's diagonal a term scaled by the prior uncertainty, only during model fitting; an adaptive version applies it only when the surrogate is confident, and the benchmarks include a 16-dimensional plasma-medicine controller. Pukdee et al. (2026) give conditions under which the Bradley-Terry model still recovers the conditional preference distribution when the data violate it, and show that the margin and the connectivity of the comparison graph govern sample efficiency. Lemma 1 of local PBO (Menn et al., 2026a) shows that gradient and Hessian estimates taken from the Laplace posterior inherit the bias of its mode; the authors add that the Hessian estimate can be especially sensitive to the lengthscale and to the conditioning of the kernel matrix, and their experiments bound the lengthscale to [0.05,0.5][0.05, 0.5] on the unit cube.

How hyperparameters are learned in practice. The approaches we found range from a weak Gamma prior with the Laplace evidence (BoTorch) and the exact skew Gaussian process evidence (Benavoli et al. 2021) to a tight log-normal prior that nearly fixes the lengthscale (Koyama et al.) and box constraints (local PBO). Whether any of them can identify lengthscales from the tens to one or two hundred comparisons a person gives is unstudied, and Section 30.3.3 states the gap and the simulation that would close it.

Sources cited in Section 27.6 3
  1. Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
  2. Pukdee et al. (2026) What Does Preference Learning Recover from Pairwise Comparison Data?
  3. Menn et al. (2026a) Local Preferential Bayesian Optimization

27.7 Settled, contested, missing #

Research status Settled, contested, missing

Settled. Under a probit preference likelihood and a Gaussian process prior, the exact posterior is a skew Gaussian process (Benavoli et al., 2021c). The original model of González et al. used a logistic link and did not state its inference method. The preference models in BoTorch and optuna-dashboard use lengthscale priors that do not scale with dimension. The Laplace likelihood Hessian is rank deficient whenever the comparison graph is disconnected, a fact of linear algebra; in practice it shows up as poor conditioning rather than singularity (inference).

Contested. Whether the error of Gaussian approximations at human noise levels is large enough to change optimization outcomes: the evidence of Benavoli et al. and Takeno et al. comes from low noise or specific settings, and they disagree about which settings are realistic. Whether the ill-conditioning that the 2026 preprint calls structural causes measurable loss on human data. Whether intransitivity is common: Chau et al.'s data are not design preferences, and a linear model did better on the NFL data.

Missing. A comparison of link functions on human data, and robustness to a wrong link. An observation model with an abstain outcome. A model of drifting utility. Evidence on whether lengthscales are identifiable at human-scale budgets. A test of dimension-scaled priors with a pairwise likelihood. A closed-loop human experiment comparing Laplace, expectation propagation, exact skew Gaussian process sampling, and the hallucination believer. A prior-fitted network trained on pairwise data. Beyond the probit link with constant noise, models of how people answer remain an open design space of unvalidated defaults (inference).

Sources cited in Section 27.7 1
  1. Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes

Further reading #

References

  1. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §27.1
  2. Astudillo, R., Li, K., Tucker, M., Cheng, C. X., Ames, A. D., and Yue, Y. (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Cited in §27.2
  3. Bemporad, A., and Piga, D. (2021). Global optimization based on active preference learning with radial basis functions. Machine Learning. Cited in §27.3
  4. Benavoli, A., and Azzimonti, D. (2024). Linearly Constrained Gaussian Processes are SkewGPs: application to Monotonic Preference Learning and Desirability. Uncertainty in Artificial Intelligence. Cited in §27.3
  5. Benavoli, A., and Azzimonti, D. (2026a). A tutorial on learning from preferences and choices with Gaussian Processes. Foundations and Trends in Machine Learning 19(1):1-120. Cited in §27.2
  6. Benavoli, A., Azzimonti, D., and Piga, D. (2020). Skew Gaussian processes for classification. Machine Learning. Cited in §27.3
  7. Benavoli, A., Azzimonti, D., and Piga, D. (2021a). A unified framework for closed-form nonparametric regression, classification, preference and mixed problems with Skew Gaussian Processes. Machine Learning. Cited in §27.3
  8. Benavoli, A., Azzimonti, D., and Piga, D. (2021c). Preferential Bayesian optimisation with skew gaussian processes. Proceedings of the Genetic and Evolutionary Computation Conference Companion. Cited in §27.2 §27.4 §27.7
  9. Benavoli, A., Azzimonti, D., and Piga, D. (2023). Learning Choice Functions with Gaussian Processes. Uncertainty in Artificial Intelligence. Cited in §27.2
  10. Bıyık, E., Palan, M., Landolfi, N. C., Losey, D. P., and Sadigh, D. (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Cited in §27.2
  11. Bıyık, E., Huynh, N., Kochenderfer, M. J., and Sadigh, D. (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. Cited in §27.3 §27.4
  12. Brochu, E., de Freitas, N., and Ghosh, A. (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Cited in §27.1
  13. Chau, S. L., González, J., and Sejdinovic, D. (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. Cited in §27.2
  14. Chu, W., and Ghahramani, Z. (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Cited in §27.1
  15. Dao, L. A., Maccarini, M., Nicora, M. L., Falerni, M. M., Mondellini, M., Veerappan, P., … Roveda, L. (2025). Experience in Engineering Complex Systems: Active Preference Learning With Multiple Outcomes and Certainty Levels. IEEE Transactions on Human-Machine Systems. Cited in §27.2
  16. Dubey, M., De Peuter, S., Wang, W., and Kaski, S. (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. preprint Cited in §27.2
  17. Erarslan, A., Sevilla Salcedo, C., Tanskanen, V., Nisov, A., Päiväkumpu, E., Aisala, H., … Mikkola, P. (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Cited in §27.2
  18. Fauvel, T., and Chalk, M. (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. preprint Cited in §27.2
  19. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §27.1
  20. Granley, J., Fauvel, T., Chalk, M., and Beyeler, M. (2023). Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. NeurIPS 2023. Cited in §27.3
  21. Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M. (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Cited in §27.1
  22. Houlsby, N., Huszár, F., Ghahramani, Z., and Hernández-lobato, J. (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. Cited in §27.2
  23. Hvarfner, C., Hellsten, E. O., and Nardi, L. (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. Cited in §27.5
  24. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §27.1
  25. Koyama, Y., Sato, I., Sakamoto, D., and Igarashi, T. (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Cited in §27.2
  26. Koyama, Y., Sato, I., and Goto, M. (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Cited in §27.1 §27.2 §27.4
  27. Kuss, M., and Rasmussen, C. E. (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. Cited in §27.4
  28. Leenders, N., Quadt, T., Cule, B., Lindelauf, R., Monsuur, H., van Oijen, J., and Voskuijl, M. (2025). DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization. arXiv. preprint Cited in §27.3
  29. Li, K., Tucker, M., Bıyık, E., Novoseller, E., Burdick, J. W., Sui, Y., … Ames, A. D. (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Cited in §27.2
  30. Li, S., Zhang, Y., Ren, Z., Liang, C., Li, N., and Shah, J. A. (2024a). Enhancing Preference-based Linear Bandits via Human Response Time. Advances in Neural Information Processing Systems. Cited in §27.2
  31. Li, Z., Liao, Y.-C., and Holz, C. (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Cited in §27.3
  32. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §27.3
  33. Menn, J., Kober, M., Brunzema, P., Stenger, D., and Trimpe, S. (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Cited in §27.6
  34. Menn, J., Stenger, D., and Trimpe, S. (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Cited in §27.2
  35. Meta Platforms, Inc. (2026e). BoTorch CHANGELOG. GitHub. software Cited in §27.5
  36. Meta Platforms, Inc. (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. software Cited in §27.1 §27.5
  37. Meta Platforms, Inc. (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Cited in §27.5
  38. Mikkola, P., Todorović, M., Järvi, J., Rinke, P., and Kaski, S. (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §27.4
  39. Nguyen, Q. P., Tay, S., Low, B. K. H., and Jaillet, P. (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Cited in §27.1 §27.2
  40. Optuna developers (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. software Cited in §27.4 §27.5
  41. Peng, S., Chen, H., and Driggs-Campbell, K. (2025). Towards Uncertainty Unification: A Case Study for Preference Learning. RSS 2025. Cited in §27.2
  42. Previtali, D., Mazzoleni, M., Ferramosca, A., and Previdi, F. (2023). GLISp-r: a preference-based optimization algorithm with convergence guarantees. Computational Optimization and Applications. Cited in §27.3
  43. Pukdee, R., Balcan, M.-F., and Ravikumar, P. (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. Cited in §27.6
  44. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §27.6
  45. Shvartsman, M., Letham, B., Bakshy, E., and Keeley, S. (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Cited in §27.1 §27.2
  46. Siivola, E., Dhaka, A. K., Andersen, M. R., González, J., García Moreno, P., and Vehtari, A. (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Cited in §27.2
  47. Simpson, E., and Gurevych, I. (2020). Scalable Bayesian preference learning for crowds. Machine Learning. Cited in §27.2
  48. Sinaga, M. A., Martinelli, J., and Kaski, S. (2026). Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization. Symposium on Probabilistic Machine Learning (ProbML 2026), Proceedings Track. Cited in §27.2
  49. Takeno, S., Nomura, M., and Karasuyama, M. (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §27.4
  50. Tatsukawa, Y., Shen, I.-C., Dogan, M. D., Qi, A., Koyama, Y., Shamir, A., and Igarashi, T. (2025). FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization. CHI 2025. Cited in §27.3
  51. Verma, A., Dai, Z., Lin, X., Jaillet, P., and Low, B. K. H. (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. Cited in §27.3
  52. Wang, H., Branke, J., and Poloczek, M. (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. Cited in §27.3
  53. Wu, K., Sanders, C., Letham, B., and Guan, P. (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. preprint Cited in §27.2
  54. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §27.1
  55. Zhang, X. (2025). PABBO code repository: evaluation config evaluate.yaml. GitHub. software Cited in §27.3
  56. Zhang, X., Huang, D., Kaski, S., and Martinelli, J. (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §27.3
  57. Zhang, R., Zhu, X., Pourebadi Khotbehsara, M., Dao, W., Bıyık, E., and Culbertson, H. (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Cited in §27.2
  58. Zhu, M., Piga, D., and Bemporad, A. (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Cited in §27.2