A Decade of Preferential Bayesian Optimization
Part IV taught preferential Bayesian optimization the way it is usually presented: a Gaussian process utility learned from comparisons (Chapter 18), a rule for choosing the next pair (Chapter 19), the forms a query can take (Section 20.1), and the bandit view of learning from duels (Chapter 21). That presentation is a snapshot. Each of its pieces arrived at a particular time, from a particular community, in answer to a particular question, and several of them have since been questioned.
This part of the book reports what research since 2017 has established, contested, or left open. This first chapter is its map. It tells the story in order, from the models that existed before the field had a name to the preprints of 2026, so that the chapters after it can each follow one thread in depth: the observation model (Chapter 27), the acquisition function (Chapter 28), the theory (Chapter 29), high dimensions (Chapter 30), and software and evaluation (Chapter 31).
The story has a surprising shape. The pipeline most people run in 2026, a Gaussian process prior with a probit link and the Laplace approximation, is a model from 2005. What changed over the decade is less the machinery than the question researchers asked of it.
The figure shows the milestones this chapter discusses and two further BoTorch releases (0.10.0 and 0.18), not every paper. A few things to look for:
- Where the theory comes from. Before 2021, every guarantee in the theory lane comes from the dueling-bandit community; kernelized results for continuous domains begin in 2021 and cluster from 2024 on.
- When people enter. Before 2022 the people lane holds only the two exoskeleton papers; afterwards it fills, with nothing in 2024.
- What is not yet peer reviewed. The hollow circles are all in 2026: the newest critiques of the default pipeline are preprints.
- Replay the decade. Turn on Hide later years, select the 2005 event, and press Next repeatedly to watch the field accumulate.
Sources cited in the introduction 4
- Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
- Meta Platforms, Inc. (2026l) CHANGELOG (versions 1.2 to 1.3)
- Facebook, Inc. (2022) ax-platform 0.2.6
- Optuna developers (2026b) optuna-dashboard 0.21.0
26.1 Before 2017 #
Three foundations existed before the field had a name, each in a different community.
The first is a model. Chu and Ghahramani (2005) placed a Gaussian process prior on a latent utility, a function that scores how much a person likes each option, and linked it to comparisons through a probit likelihood: the probability that is preferred to is the standard normal distribution function applied to the scaled utility difference (Section 16.3). Because that likelihood is not Gaussian, the posterior has no closed form, and they replaced it with the Laplace approximation, a Gaussian centered at the posterior's peak (Section 17.2). This is the model of Chapter 18, and it is still the default two decades later.
The second is an interactive system. Brochu et al. (2007) used active preference learning, in which the system itself decides which options to show next, to help people design materials for computer graphics by choosing from a gallery of candidates, with the loop of Section 19.5: show options, record a choice, update the model, choose what to show next.
The third is a problem formulation from information retrieval. Yue and Joachims (2009) cast the interactive optimization of a retrieval system, such as a search engine learning from its users, as a dueling bandit problem: a learner repeatedly picks two options and observes only which one wins (Section 21.1). The framing brought the analytic tools of online learning, regret bounds above all, to learning from comparisons.
The three lived apart: Gaussian process preference learning in machine learning, galleries in graphics and interaction design, dueling bandits in online learning. Much of the decade that followed can be read as these communities meeting, slowly and incompletely.
Sources cited in Section 26.1 3
- Chu and Ghahramani (2005) Preference learning with Gaussian processes
- Brochu et al. (2007) Active Preference Learning with Discrete Choice Data
- Yue and Joachims (2009) Interactively optimizing information retrieval systems as a dueling bandits problem
26.2 2017 to 2019: how to ask #
The name. González et al. (2017) defined the problem on the dueling space, the set of all pairs of inputs, gave it the name preferential Bayesian optimization, and proposed three acquisition functions: pure exploration, Copeland expected improvement, and dueling Thompson sampling (Section 19.2). It came without convergence theory, that is, without a statement of how fast the recommended option approaches the best one as queries accumulate.
Guarantees on the bandit side. The dueling-bandit community already had formal results for closely related problems. The quantity those results control is regret: the utility lost, summed over all queries, by showing the options the algorithm chose instead of the best one (Chapter 13). SelfSparring, a method for duels among several options at once, was proved to converge asymptotically (Sui et al., 2017b); Kumagai (2017) gave a regret bound for dueling bandits over a continuous space, in the same year the name appeared; and StageOpt handled unknown safety constraints, with theorems stated for numerical observations and a preference variant, without a convergence theorem of its own, applied to spinal cord stimulation (Sui et al., 2018b) (Section 28.7).
One sometimes reads that learning from comparisons had no formal guarantees before the kernelized regret bounds of 2024. The claim is right about one thing: the Gaussian process formulation of González et al. came without a rate, which a 2018 survey of dueling bandits pointed out, calling it "a pure Bayesian optimization approach without theoretical guarantees on convergence rate" (Sui et al., 2018a). It is wrong as a statement about the field. SelfSparring's asymptotic convergence (Sui et al., 2017b) and Kumagai's continuous-domain regret bound (Kumagai, 2017) both predate 2019, as does a clinical application of preference feedback, StageOpt's (Sui et al., 2018b). The accurate statement is narrower: guarantees existed for bandit formulations of the problem, while the Gaussian process pipeline that practitioners run had none, and as Chapter 29 reports, the analyzed algorithms and the practiced pipeline are still not the same thing.
Changing the question's shape. In graphics, Koyama et al. (2017) took a different route. Their sequential line search turns every query into a single slider: a crowd worker drags along a line through the design space and stops at the point they like best (Section 20.2). This paper began a line of human-computer interaction research in which the form of the query, rather than the rule for choosing it, is the main variable.
The question of the period. The central question of these years was how to ask a person. Machine learning improved the rules for choosing comparisons, while human-computer interaction mainly changed the shape of the query and rarely the acquisition function (inference). The figure shows no milestone in 2019; work continued, for example on letting people answer "about the same" (Bıyık et al., 2019) (Section 27.2).
Sources cited in Section 26.2 7
- González et al. (2017) Preferential Bayesian Optimization
- Sui et al. (2017b) Multi-dueling Bandits with Dependent Arms
- Kumagai (2017) Regret Analysis for Continuous Dueling Bandit
- Sui et al. (2018b) Stagewise Safe Bayesian Optimization with Gaussian Processes
- Sui et al. (2018a) Advancements in Dueling Bandits
- Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
26.3 2020 to 2021: tools and inference settle #
A default implementation. In April 2020, version 0.2.3 of BoTorch, Meta's
open-source Bayesian optimization library built on PyTorch (Section 14.8),
added PairwiseGP, a model for pairwise comparison data
(Meta Platforms, Inc., 2026e). It implements the probit link with the Laplace
approximation, and from then on that combination has been the most widely used
implementation of preferential Bayesian optimization: from 2020 the field had a
de facto standard, and that standard was the 2005 model (Section 27.5).
Queries in a subspace. The same year brought methods that handle more dimensions by restricting each query to a low-dimensional subspace: in projective preferential Bayesian optimization the person chooses the best point along a line through the space (Mikkola et al., 2020), and the Sequential Gallery shows a two-dimensional plane of designs as a grid (Koyama et al., 2020) (Section 20.3). In robotics, CoSpar learned exoskeleton walking gaits from a wearer's preferences (Tucker et al., 2020b), and LineCoSpar extended it to six gait parameters, tested with six able-bodied participants (Tucker et al., 2020a); these two papers began the line of exoskeleton applications that Section 33.1 follows and Chapter 24 works through.
The exact posterior. In 2021, Benavoli et al. (2021c) proved that the exact posterior over the utility under a probit preference likelihood is a skew Gaussian process: a distribution over functions like a Gaussian process, except that its marginals are skewed rather than symmetric (Section 17.6). A Gaussian approximation such as Laplace's cannot represent that skew, and from this point the quality of posterior inference became a point of dispute (Section 27.4).
The first kernelized bound, and an alternative. Also in 2021, Kirschner and Krause (2021) gave the first cumulative regret bound for kernelized dueling feedback, in which the unknown utility is assumed smooth in the sense defined by a kernel, so the result covers continuous domains (Section 21.3); their feedback model is the utility difference plus noise, not the probit or logistic link of the preference models. Control engineering attacked the same problem independently: GLISp fits a radial basis function surrogate, a weighted sum of bumps, to the observed preferences, with no probability model at all (Bemporad and Piga, 2021) (Section 27.3).
By the end of 2021, then, the field had a standard implementation, a known weakness in that implementation's inference, and the beginning of a theory for continuous domains. The acquisition functions in use were still largely heuristics.
Sources cited in Section 26.3 8
- Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Tucker et al. (2020b) Preference-Based Learning for Exoskeleton Gait Optimization
- Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- Kirschner and Krause (2021) Bias-Robust Bayesian Optimization via Dueling Bandits
- Bemporad and Piga (2021) Global optimization based on active preference learning with radial basis functions
26.4 2022 to 2023: the decision-theoretic turn #
EUBO. The heuristics gave way to a principled rule. In Bayesian optimization with preference exploration (BOPE), a person's preferences over the outcomes of an experiment are learned from comparisons while the experiments run. For that setting, Lin et al. (2022) proposed the expected utility of the best option, EUBO (Section 19.4), and proved it one-step Bayes optimal: if the session ended after one more answer, no other query would yield a better expected final recommendation (Section 19.4.1). Astudillo et al. (2023) generalized it to qEUBO, for queries that show several options at once and for answers with logistic noise, and proved that an adapted batch version of expected improvement is not asymptotically consistent. This was the decision-theoretic turn, and qEUBO became the basis of BoTorch's preference acquisition functions; BoTorch, Ax, and optuna-dashboard shipped the new rules within months (Section 31.1).
Inference, measured. Takeno et al. (2023) measured how far the Laplace approximation and expectation propagation are from the exact skew posterior, and proposed the hallucination believer: draw one sample of the latent utilities from the posterior, treat it as measured data, and apply any standard acquisition function (Section 19.3). The method became the basis of the preferential sampler in optuna-dashboard, a web interface for the Optuna optimization library (Optuna developers, 2026b).
People become the question. At the same time, human-computer interaction began to measure what happens to the person in the loop. Novice designers whose design process was led by a multi-objective optimizer reported lower agency and ownership than those who led it themselves, in a study that used performance objectives, not comparisons (Chan et al., 2022) (Section 32.2); in a three-month field deployment, most rating loops never reached the optimization stage or did not converge (Ou et al., 2022) (Section 31.7); and experts iterated more than novices and ended less satisfied (Ou et al., 2023) (Section 32.6).
A neighbor grows large. In 2023, direct preference optimization (DPO) put the Bradley-Terry likelihood, under which the probability that one option beats another is the logistic function of their utility difference (Section 16.4), at the center of aligning large language models with human preferences, fitting the language model to human comparisons directly rather than through a reward model as reinforcement learning from human feedback (RLHF) does (Rafailov et al., 2023). The same year, the ICML workshop The Many Facets of Preference-Based Learning brought dueling bandits, RLHF, social choice, and optimization together (ICML, 2023). Chapter 35 follows the traffic between the two fields.
Sources cited in Section 26.4 9
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Optuna developers (2026b) optuna-dashboard 0.21.0
- Chan et al. (2022) Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Ou et al. (2023) The Impact of Expertise in the Loop for Exploring Machine Rationality
- Rafailov et al. (2023) Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- ICML (2023) The Many Facets of Preference-Based Learning
26.5 2024 to 2026: noise, theory, and scrutiny #
The last three years divide into two stages. In 2024 and 2025, theory, high-dimensional diagnosis, language models, and questions of legitimacy all advanced at once. In 2026, the work concentrated on flaws in the default pipeline and on simpler alternatives to it.
26.5.1 2024 and 2025: theory, dimension, language models, legitimacy #
Bounds under the Bradley-Terry link. Kernelized regret upper bounds, guarantees that cumulative regret grows no faster than a stated rate, arrived for the Bradley-Terry link with POP-BO (Xu et al., 2024b), MaxMinLCB (Pásztor et al., 2024), and MR-LPF, whose upper bound is of the same order as for scalar feedback (Kayal et al., 2025). Matching upper bounds do not show that a comparison carries as much information as a number; they show only that the guarantees are of the same order. Section 29.4 states each rate with its assumptions.
An amortized optimizer. PABBO trains a neural network in advance, on many synthetic tasks, to propose the next pair directly, so that a query costs one forward pass instead of fitting a model and optimizing an acquisition function (Zhang et al., 2025a). It is the only amortized optimizer for pairwise preferences (Section 27.3).
High dimensions, re-diagnosed. In ordinary Bayesian optimization with
numerical observations, a run of papers attributed the familiar failure in high
dimensions to the prior and the initialization rather than to the method itself
(Hvarfner et al., 2024; Xu et al., 2025b; Papenmeier et al., 2025b). The
remedy is a dimension-scaled prior, a prior on the kernel lengthscale whose
typical value grows with the number of inputs (Section 9.5). In
September 2024, BoTorch 0.12.0 switched most of its models to such priors but
explicitly excluded PairwiseGP (Meta Platforms, Inc., 2026e), a gap that
Section 30.1 examines.
Language models and legitimacy. Large language models entered conversational preference elicitation (Austin et al., 2024a) and the active collection of alignment data (Dwaracherla et al., 2024); human-computer interaction turned to population priors, which transfer what earlier users preferred to a new user (Li et al., 2025a), and to collaboration in natural language between designer and optimizer (Niwa et al., 2025). Alignment research also formalized a worry that preferential optimization shares: a system that learns from preferences may change the preferences it measures. Carroll et al. (2024) compared eight notions of alignment for preferences that can change and found that each either rewards the system for undue influence on the person or is overly risk-averse, and Williams et al. (2025) found that learners optimized on user feedback learn to target the users most open to influence (Section 40.3).
26.5.2 2026: the default pipeline under scrutiny #
Theory. Thompson sampling with preference feedback, PF-TS, has an upper bound of order for a fully sequential algorithm (Lazzaro et al., 2026), against MR-LPF's for a batched one with a finite candidate set and a warm-up period (Kayal et al., 2025); here is the maximum information gain of the kernel, which grows slowly for smooth kernels (Section 6.5).
Flaws in the default. Two 2026 preprints examined the pipeline of
PairwiseGP, the Laplace approximation, and EUBO: EUBO's queries collapse
toward the estimated best option (Wu and Gardner, 2026), and its pairs, sharing no
input with earlier queries, make the Laplace likelihood Hessian rank deficient
(Shao et al., 2026). Section 19.6 introduced both. The observation
model was extended to experiments that fail (Menn et al., 2026b), and local
preferential Bayesian optimization, also a preprint, took the method to about
100 dimensions (Menn et al., 2026a) (Section 30.4).
Simpler alternatives and people. In ordinary high-dimensional Bayesian optimization, a spherical mapping of the inputs combined with Bayesian linear regression (Section 5.4) reached the state of the art on tasks with 60 to 6000 dimensions (Doumont et al., 2026). Ax added preference optimization and trials in which a language model turns free-text feedback into comparisons (Meta Platforms, Inc., 2026l; Kobalczyk et al., 2026) (Section 35.2). And the human studies of 2026 that used strong comparison conditions mostly found null results, or found that people tuning by hand did about as well: cost-aware Bayesian optimization for prototyping interactive devices reached the same performance at about 67% of the cost, with no difference in final quality (Langerak et al., 2026); priors learned from user models helped only at the second and third iterations, in a study with 12 participants and a performance objective (Liao et al., 2026); and 11 healthy adults who tuned their own exoskeleton assistance with a thumbstick remote control reached a 16.6% metabolic reduction in about 10.9 minutes, comparable to what algorithmic tuning reports (Schäfer et al., 2026). Simpler methods, including people tuning by hand, have thereby become controls that a new method must be compared against (Section 32.8). The tutorial of Benavoli and Azzimonti on learning from preferences and choices with Gaussian processes was also formally published (Benavoli and Azzimonti, 2026a), though there is still no dedicated survey of the field (Section 31.8.2).
Sources cited in Section 26.5 26
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
- Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Hvarfner et al. (2024) Vanilla Bayesian Optimization Performs Great in High Dimensions
- Xu et al. (2025b) Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization
- Papenmeier et al. (2025b) Understanding High-Dimensional Bayesian Optimization
- Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
- Austin et al. (2024a) Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation
- Dwaracherla et al. (2024) Efficient Exploration for LLMs
- Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Carroll et al. (2024) AI Alignment with Changing and Influenceable Reward Functions
- Williams et al. (2025) On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
- Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
- Wu and Gardner (2026) Knowledge Gradient for Preference Learning
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Menn et al. (2026b) Preferential Bayesian Optimization with Crash Feedback
- Menn et al. (2026a) Local Preferential Bayesian Optimization
- Doumont et al. (2026) We Still Don't Understand High-Dimensional Bayesian Optimization
- Meta Platforms, Inc. (2026l) CHANGELOG (versions 1.2 to 1.3)
- Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback
- Langerak et al. (2026) Cost-Aware Bayesian Optimization for Prototyping Interactive Devices
- Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Benavoli and Azzimonti (2026a) A tutorial on learning from preferences and choices with Gaussian Processes
26.6 How the central question moved #
Read as a sequence of questions, the decade turned three times (inference).
From 2017 to 2021 the question was how to ask a person, and how to infer from
the answers. The dueling formulation, sequential line search, galleries and
projections, PairwiseGP, and the skew Gaussian process all answer some part of
it.
From 2022 to 2025 it was whether the acquisition function is principled, and whether it comes with guarantees. EUBO and qEUBO answered the first part with one-step Bayes optimality; POP-BO, MaxMinLCB, and MR-LPF answered the second with regret bounds under the Bradley-Terry link.
From 2025 to 2026 it became four questions at once: is the observation model right, do people answer the way the model assumes, are simpler methods already enough, and does the system change the preferences it measures? The collapse of EUBO, the rank-deficient Hessian, the null results under strong controls, the linear model that wins in high dimension, and the alignment results on influence all belong here. This last turn is the book's thesis seen from the literature: the bottleneck moved from algorithms to measurement (Section 45.1).
The periods overlap in 2025, when the second question was still being answered and the third was already being asked, which is why Figure 26.1 shows both for that year.
The default pipeline of 2026, a Gaussian process prior with a probit link and the Laplace approximation, is the model of 2005. What changed over the decade is the question asked of it: first how to ask, then whether the choice of query is principled and guaranteed, and finally whether the model of the person is right at all.
Each of the following chapters takes up one of these threads. Chapter 27 asks whether the model of a human answer is right; Chapter 28, whether the query rules hold up; Chapter 29, what is actually proved; Chapter 30, where the method stops working and why; and Chapter 31, what the software does by default and how methods are compared. Part VII asks the human questions, and Part IX asks what a preference is in the first place.
26.7 Publication and community #
The volume of work grew after 2023, but from a small base: on arXiv, entries whose abstracts contain preferential, Bayesian, and optimization (or optimisation) numbered 4 in 2023, 5 in 2024, 13 in 2025, and 12 in the first nine months of 2026 (arXiv, 2026b), a count that misses papers using other words, such as "dueling" or "human feedback" (inference). The research is spread across the venues of machine learning, control, robotics, and human-computer interaction, and these communities cite one another little: the 2021 survey of dueling bandits in the Journal of Machine Learning Research does not cite González et al. 2017 (Bengs et al., 2021). Table 31.6 maps the main groups by discipline; the groups that prove regret bounds do not run human studies, and the groups that run human studies rarely change the acquisition function (inference). Section 31.8 gives the fuller account.
Sources cited in Section 26.7 2
- arXiv (2026b) Abstract search: preferential AND Bayesian AND (optimization OR optimisation)
- Bengs et al. (2021) Preference-based Online Learning with Dueling Bandits: A Survey
26.8 Settled, contested, missing #
Settled. The default pipeline did not change over the decade: the Gaussian
process, probit, and Laplace model of Chu and Ghahramani (2005), implemented in
BoTorch's PairwiseGP since April 2020 (Meta Platforms, Inc., 2026e). EUBO and qEUBO
are one-step Bayes optimal with noise-free answers
(Lin et al., 2022; Astudillo et al., 2023). Formal guarantees for learning from
comparisons existed on the dueling-bandit side well before 2024
(Sui et al., 2017b; Kumagai, 2017), while kernelized bounds
under the Bradley-Terry link date from 2024 (Xu et al., 2024b). BoTorch's
switch to dimension-scaled priors in 2024 explicitly excluded the preference
model (Meta Platforms, Inc., 2026e).
Contested. Whether the flaws reported in 2026, EUBO's collapse toward the current best and the rank-deficient Hessian, cost anything on human tasks; both rest on preprints (Wu and Gardner, 2026; Shao et al., 2026). Whether preferential optimization beats simpler alternatives once the comparison is strong: the controlled human studies of 2026 found modest or no advantages (Langerak et al., 2026; Liao et al., 2026; Schäfer et al., 2026), and in scalar high dimension a linear model reached the state of the art (Doumont et al., 2026).
Missing. A dedicated survey of preferential Bayesian optimization. Exchange between the communities: the 2021 JMLR survey of dueling bandits does not cite the paper that named the field (Bengs et al., 2021). A study that randomizes people to different acquisition functions with the same interface and budget, and a comparison with expert manual tuning on a preregistered endpoint (Section 47.4).
Sources cited in Section 26.8 14
- Chu and Ghahramani (2005) Preference learning with Gaussian processes
- Meta Platforms, Inc. (2026e) BoTorch CHANGELOG
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Sui et al. (2017b) Multi-dueling Bandits with Dependent Arms
- Kumagai (2017) Regret Analysis for Continuous Dueling Bandit
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Wu and Gardner (2026) Knowledge Gradient for Preference Learning
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Langerak et al. (2026) Cost-Aware Bayesian Optimization for Prototyping Interactive Devices
- Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Doumont et al. (2026) We Still Don't Understand High-Dimensional Bayesian Optimization
- Bengs et al. (2021) Preference-based Online Learning with Dueling Bandits: A Survey
Further reading #
- González et al. (2017) is the paper that named the field; read it next to Chu and Ghahramani (2005), whose model the field went on to use instead of the dueling formulation.
- Sui et al. (2018a) and Bengs et al. (2021) are two surveys of dueling bandits; together they show what the bandit community knew and what it did not cite.
- Lin et al. (2022) and Astudillo et al. (2023) mark the decision-theoretic turn.
- Benavoli and Azzimonti (2026a) is the most complete single reference on Gaussian process models for preferences and choices.
- The BoTorch changelog (Meta Platforms, Inc., 2026e) is a compact history of what practitioners could actually run, release by release.
References
- (2026b). Abstract search: preferential AND Bayesian AND (optimization OR optimisation). arXiv API. non-peer-reviewed Cited in §26.7
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §26.4 §26.8
- (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Cited in §26.5
- (2021). Global optimization based on active preference learning with radial basis functions. Machine Learning. Cited in §26.3
- (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. Cited in §26.7 §26.8
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Cited in §26.2
- (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Cited in §26.1
- (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. Cited in §26.5
- (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. Cited in §26.4
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Cited in §26.1 §26.8
- (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). Cited in §26.5 §26.8
- (2024). Efficient Exploration for LLMs. ICML 2024. Cited in §26.5
- (2022). ax-platform 0.2.6. PyPI. software
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §26.2
- (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. Cited in §26.5
- (2023). The Many Facets of Preference-Based Learning. ICML 2023 workshop page. non-peer-reviewed Cited in §26.4
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §26.5
- (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. Cited in §26.3
- (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Cited in §26.5
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Cited in §26.2
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Cited in §26.3
- (2017). Regret Analysis for Continuous Dueling Bandit. Advances in Neural Information Processing Systems. Cited in §26.2 §26.8
- (2026). Cost-Aware Bayesian Optimization for Prototyping Interactive Devices. CHI 2026. Cited in §26.5 §26.8
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §26.5
- (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Cited in §26.5
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Cited in §26.5 §26.8
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §26.4 §26.8
- (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Cited in §26.5
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Cited in §26.5
- (2026e). BoTorch CHANGELOG. GitHub. software Cited in §26.3 §26.5 §26.8
- (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Cited in §26.5
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §26.3
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Cited in §26.5
- (2026b). optuna-dashboard 0.21.0. PyPI. software Cited in §26.4
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §26.4
- (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. Cited in §26.4
- (2025b). Understanding High-Dimensional Bayesian Optimization. ICML 2025, PMLR 267:47902-47923. Cited in §26.5
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Cited in §26.5
- (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. Cited in §26.4
- (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §26.5 §26.8
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §26.5 §26.8
- (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. Cited in §26.2 §26.8
- (2018a). Advancements in Dueling Bandits. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. doi:10.24963/ijcai.2018/776. Cited in §26.2
- (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. Cited in §26.2
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §26.4
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Cited in §26.3
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Cited in §26.3
- (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Cited in §26.5
- (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Cited in §26.5 §26.8
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §26.5 §26.8
- (2025b). Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization. ICLR 2025 (oral). Cited in §26.5
- (2009). Interactively optimizing information retrieval systems as a dueling bandits problem. Proceedings of the 26th Annual International Conference on Machine Learning. Cited in §26.1
- (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §26.5