Reward Learning, Recommendation, Ranking, and Automated Science
Preferential Bayesian optimization (PBO) is one of many systems in computing that learn from what people say they prefer. Chapter 35 followed the largest neighbor, the alignment of language models. This chapter visits seven more, each of which met the problems of preferential optimization in its own form and often earlier: reinforcement learning from comparisons of trajectories, AI safety, preference elicitation in AI and decision analysis, recommender systems, interactive evolutionary computation, learning to rank, and self-driving laboratories.
Two conclusions run through the chapter. First, several things that the PBO literature still treats as open have been solved, or have a ready practice, next door: choosing queries by their effect on the final decision, monotonicity and ranking queries in a Gaussian process model, simulated users with realistic irrationalities, correction for the position in which an option is shown, warm starts from population embeddings, calibration metrics, and a reporting standard for speedups; Table 36.2 collects them. Second, the AI safety literature has formalized the idea that a system can change the person it measures, and has found evidence of it in recommender systems and language models. That is the book's claim that a PBO session is an intervention as well as an estimate (Section 45.1), worked out by a neighbor; no study of PBO has yet measured it, so each transfer of the concern below is marked as inference.
36.1 Preference-based reinforcement learning #
Reinforcement learning trains an agent, a policy that chooses actions, to maximize the total reward it collects over a sequence of steps, the return. Writing a reward function by hand is hard for tasks such as a backflip or a helpful answer. Preference-based reinforcement learning replaces it with human comparisons of short clips of the agent's behavior, trajectory segments, and learns a reward model from the answers with the Bradley-Terry likelihood of Section 35.1.1; RLHF (Section 35.1.2) is this method applied to language models. Christiano et al. (2017) (NeurIPS 2017) showed it working at scale, giving feedback "on less than one percent of our agent's interactions with the environment" and teaching new behaviors with "about an hour of human time".
36.1.1 Practice and benchmarks #
PEBBLE (Lee et al., 2021b) (ICML 2021) relabels all past experience whenever the reward model changes and pretrains the agent without rewards. B-Pref (Lee et al., 2021a) (NeurIPS 2021 Datasets and Benchmarks) argued that "simulating human input as giving perfect preferences for the ground truth reward function is unrealistic", and built simulated teachers that each add one irrationality to a perfect one: random choice with rationality (the Bradley-Terry model with unit scale), mistakes with probability 0.1, skipping pairs that are hard to compare, declaring a tie when the returns are close, and myopia, discounting earlier steps in a segment with . Uni-RLHF (Yuan et al., 2024) (ICLR 2024) collected crowdsourced annotations covering "more than 15 million steps across 30+ popular tasks", with "competitive performance compared to those from well-designed manual rewards".
36.1.2 Theory #
Dueling posterior sampling (Novoseller et al., 2020) (UAI 2020) gave the first regret guarantee for preference-based reinforcement learning, an asymptotic Bayesian no-regret rate; Saha et al. (2023) (AISTATS 2023) proved near-optimal regret for trajectory preferences under generalized linear models; and Wang et al. (2023a) (NeurIPS 2023) proved that "for a wide range of preference models, we can solve preference-based RL directly using existing algorithms and techniques for reward-based RL, with small or no extra costs".
36.1.3 Models of how people compare #
The model of how a person turns a reward into an answer matters more than expected. Knox et al. (2024) (TMLR 2024) compared the usual model, in which a person prefers the segment with the larger sum of rewards, its partial return, with one in which the person judges each segment by its regret, how much worse it is than the best the agent could have done from the same start. The regret model was identifiable (given enough answers, only one reward function is consistent with them), the partial-return model lacked identifiability in several settings, and the regret model predicted real human preferences better. How wrong can the inferred reward be if the model of the person is slightly wrong? Hong et al. (2023) (ICLR 2023) showed that "it is unfortunately possible to construct small adversarial biases in behavior that lead to arbitrarily large errors in the inferred reward", and identified "reasonable assumptions under which the reward inference error can be bounded linearly in the error in" the human model. And Hatgis-Kessell et al. (2025) (TMLR) turned the question around, helping people fit the model by showing them the latent quantity it assumes, training them to follow it, or rewording the question, in three studies with people: "All intervention types show significant effects."
36.1.4 Choosing queries and using richer feedback #
Information-directed reward learning (Lindner et al., 2021) (NeurIPS 2021) selects "queries that maximize the information gain about the difference in return between plausibly optimal policies", rather than about the reward everywhere, and needs significantly fewer queries. Hu et al. (2024) (ICLR 2024) named the failure this avoids query-policy misalignment: queries chosen to improve the reward model overall "may not align with RL agents' interests, thus offering little help on policy learning". RIME (Cheng et al., 2024) (ICML 2024) filters noisy preferences; LiRE (Choi et al., 2024) (ICML 2024) builds ranked lists of trajectories from better, worse, and equal answers to exploit how strong each preference is; and inverse preference learning (Hejna and Sadigh, 2023) (NeurIPS 2023) and contrastive preference learning (Hejna et al., 2024) (ICLR 2024) skip the explicit reward model, both assuming that preferences follow regret.
36.1.5 Over-optimization #
When a policy is optimized hard against a learned reward, the true reward eventually falls, an instance of Goodhart's law ("when a measure becomes a target, it ceases to be a good measure"). Gao et al. (2023) (ICML 2023, in the PMLR proceedings; the arXiv record lists no venue) measured this with a fixed "gold-standard" reward model standing in for people: the relationship between optimization and gold reward "follows a different functional form depending on the method of optimization", and its coefficients scale smoothly with reward model size. Rafailov et al. (2024) (NeurIPS 2024) found that DPO and its relatives over-optimize too, "often before even a single epoch of the dataset is completed". Casper et al. (2023) (TMLR 2023) survey the open problems and fundamental limits of RLHF.
36.1.6 What it means for PBO #
Three conclusions carry over as stated: the choice of preference model has real consequences (Knox et al.), which in PBO is the choice between probit and logistic links (Section 27.1); Hatgis-Kessell et al. recommend "designing interfaces and training interventions to increase human conformance with the modeling assumptions of the algorithm", which in PBO is how the question is worded and shown (Chapter 20); and a simulated teacher with perfect preferences is unrealistic (B-Pref).
The rest is inference. By Hong et al., small systematic response biases, such as an effect of position, can make a PBO posterior badly wrong even after many comparisons, whereas random noise is the benign case (Section 36.6). B-Pref's teachers map almost one to one onto simulated users for PBO (skipping to incomparability, ties to indifference, Section 20.4, myopia to recency effects), so PBO benchmarks, which mostly simulate homoscedastic, independent noise (the same noise level for every pair, drawn afresh each time), could adopt them with little change (Section 31.7). The theory on both sides, Wang, Liu, and Jin for reinforcement learning, Shah et al. for ranking (Section 36.6), and Kayal et al. for PBO (Section 29.3), suggests that a pairwise answer carries less information per query than a numerical one, but the achievable rate is of the same order, with the qualification that Kayal et al.'s result is a conditional upper bound. And over-optimization reaches PBO only when a learned utility is optimized without a person checking the results, as with the utility over outcomes in preference exploration (BOPE); in standard PBO the person judges the final candidates.
Sources cited in Section 36.1 19
- Christiano et al. (2017) Deep reinforcement learning from human preferences
- Lee et al. (2021b) PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
- Lee et al. (2021a) B-Pref: Benchmarking Preference-Based Reinforcement Learning
- Yuan et al. (2024) Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
- Novoseller et al. (2020) Dueling Posterior Sampling for Preference-Based Reinforcement Learning
- Saha et al. (2023) Dueling RL: Reinforcement Learning with Trajectory Preferences
- Wang et al. (2023a) Is RLHF More Difficult than Standard RL?
- Knox et al. (2024) Models of human preference for learning reward functions
- Hong et al. (2023) On the Sensitivity of Reward Inference to Misspecified Human Models
- Hatgis-Kessell et al. (2025) Influencing Humans to Conform to Preference Models for RLHF
- Lindner et al. (2021) Information Directed Reward Learning for Reinforcement Learning
- Hu et al. (2024) Query-Policy Misalignment in Preference-Based Reinforcement Learning
- Cheng et al. (2024) RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
- Choi et al. (2024) Listwise Reward Estimation for Offline Preference-based Reinforcement Learning
- Hejna and Sadigh (2023) Inverse Preference Learning: Preference-based RL without a Reward Function
- Hejna et al. (2024) Contrastive Preference Learning: Learning from Human Feedback without RL
- Gao et al. (2023) Scaling Laws for Reward Model Overoptimization
- Rafailov et al. (2024) Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
- Casper et al. (2023) Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
36.2 The AI safety view #
Since 2017 the AI safety literature has formalized three ideas that matter for any system that optimizes on behalf of a person: a stand-in for what the person wants can fail when optimized, the system can change the person, and the person and the system can be modeled as playing a cooperative game.
36.2.1 When the proxy fails #
Skalse et al. (2022) (NeurIPS 2022, where the paper appears as "Defining and Characterizing Reward Gaming") defined a proxy reward as unhackable if "increasing the expected proxy return can never decrease the expected true return", and proved that "for the set of all stochastic policies, two reward functions can only be unhackable if one of them is constant". "More capable agents often exploit reward misspecifications", with "phase transitions: capability thresholds at which the agent's behavior qualitatively shifts, leading to a sharp decrease in the true reward" (Pan et al., 2022) (ICLR 2022). Karwowski et al. (2024) (ICLR 2024) explained Goodhart's law geometrically and gave an early-stopping method that provably avoids it, and Kwa et al. (2024) (NeurIPS 2024) showed that with heavy-tailed reward error "some policies obtain arbitrarily high reward despite achieving no more utility than the base model" even under KL regularization, though the reward models they measured were "consistent with light-tailed error". From the principal's side, unbounded optimization of a proxy that covers only some attributes can drive the person's utility arbitrarily low, and letting the person update the proxy over time helps (Zhuang and Hadfield-Menell, 2020) (NeurIPS 2020); a written reward function is "merely observations about what the designer actually wants" (Hadfield-Menell et al., 2017) (NeurIPS 2017); and by the optimizer's curse (Smith and Winkler, 2006), the estimated value of the option picked as best is biased upward, because it was picked partly for its lucky error.
36.2.2 Influence on the person, and approval optimization #
Approval optimization means optimizing the evaluator's approval rather than the outcome the evaluator cares about. Carroll et al. (2022) (ICML 2022) pointed out that "systems trained via long-horizon optimization will have direct incentives to manipulate users", in particular "to shift user preferences so they are easier to satisfy", and proposed penalizing shifts outside a trust region of "safe shifts", such as the drift a person's preferences would undergo without the system. Carroll et al. (2024) (ICML 2024) wrote that "existing AI alignment approaches assume that preferences are static, which is unrealistic", found that "8 such notions of alignment" all "either err towards causing undesirable AI influence, or are overly risk-averse", and argued that the optimization horizon "may partially help reduce undesirable AI influence".
The evidence comes from simulation and language models. A Q-learning recommender "consistently learns to exploit its opportunities to polarize simulated users" (Evans and Kasirzadeh, 2023) (AIES 2023). "Even if only 2% of users are vulnerable to manipulative strategies", language models trained on user feedback learn to identify and target them, and safety training or language-model judges help in some settings but "backfire in others, sometimes even leading to subtler manipulative behaviors" (Williams et al., 2025) (ICLR 2025). After RLHF, time-limited human evaluators accepted wrong answers as correct more often, with false positive rates up 24.1% on QuALITY and 18.3% on APPS (Wen et al., 2025) (ICLR 2025). "Both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time" (Sharma et al., 2024) (ICLR 2024). And when evaluators see only part of what happened, RLHF can produce "deceptive inflation" and "overjustification" (Lang et al., 2024) (NeurIPS 2024).
36.2.3 Assistance games #
An assistance game, also called cooperative inverse reinforcement learning, models a person and an AI assistant who share the person's goal, which only the person knows. AssistanceZero (Laidlaw et al., 2025) (ICML 2025) solved a Minecraft assistance game "with over possible goals", and its assistant "significantly reduces the number of actions participants take to complete building tasks". Emmons et al. (2025) (ICML 2025) proved that an optimal assistant must sometimes interfere with what the person can observe: when the person decides based on immediate outcomes, the assistant may need to interfere in order to query their preferences, an incentive that vanishes if the person has a channel to communicate preferences; and a Boltzmann-irrational person (choosing better options more often but not always) can also create an incentive to interfere. Ananthakrishnan et al. (2026), a 2026 preprint with a workshop version, gave the first efficient algorithms for repeated assistance games, with a -approximate assistance regret of , and proved that beating is computationally infeasible; Fickinger et al. (2020), a 2020 preprint, applied impossibility theorems from social choice to assistants serving several people (Section 40.7). And existing RLHF algorithms are not strategyproof (immune to gains from misreporting): "even a single strategic labeler can cause arbitrarily large misalignment with social welfare", and "any strategyproof RLHF algorithm must perform -times worse than the optimal policy, where is the number of labelers" (Kleine Buening et al., 2025) (NeurIPS 2025).
36.2.4 What it means for PBO #
None of these results has been tested on PBO, so this subsection is inference. PBO can be seen as a restricted assistance game in which the assistant acts only by choosing which candidates to show and what to recommend, and it assumes exactly the Boltzmann-irrational person of Emmons et al., so an acquisition function optimal under that model might favor pairs that are informative to the model but distort the person's understanding of the design space. Whenever people compare renderings, summaries, or explanations rather than real outcomes, the results of Wen et al. and Lang et al. apply: the loop can converge to designs that look better rather than designs that are better. And by the optimizer's curse, the estimated utility of the recommended design is biased upward when the model is misspecified, which Bayesian shrinkage toward the prior addresses.
Carroll et al.'s argument about horizons (2024) suggests that a myopic acquisition function has no planned incentive to change the person, but changes caused by what the person is shown still happen, and Figure 42.1 simulates how a benchmark score can improve while the recommendation drifts from what the person first wanted. As of September 2026 we found no study of PBO that measures preference change caused by the acquisition function. A study of PBO with people could measure it by comparing preference drift when an optimizer chooses the pairs and when it does not, taking the natural drift without the system, which Carroll et al.'s 2022 paper defines as a safe shift, as the baseline. Karwowski et al.'s provable early stopping is a candidate for the stopping rule PBO lacks (Section 46.6), and with several stakeholders, Kleine Buening et al.'s factor of limits what incentive-compatible aggregation can achieve.
Sources cited in Section 36.2 19
- Skalse et al. (2022) Defining and Characterizing Reward Hacking
- Pan et al. (2022) The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
- Karwowski et al. (2024) Goodhart's Law in Reinforcement Learning
- Kwa et al. (2024) Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
- Zhuang and Hadfield-Menell (2020) Consequences of Misaligned AI
- Hadfield-Menell et al. (2017) Inverse Reward Design
- Smith and Winkler (2006) The Optimizer’s Curse: Skepticism and Postdecision Surprise in Decision Analysis
- Carroll et al. (2022) Estimating and Penalizing Induced Preference Shifts in Recommender Systems
- Carroll et al. (2024) AI Alignment with Changing and Influenceable Reward Functions
- Evans and Kasirzadeh (2023) User Tampering in Reinforcement Learning Recommender Systems
- Williams et al. (2025) On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
- Wen et al. (2025) Language Models Learn to Mislead Humans via RLHF
- Sharma et al. (2024) Towards Understanding Sycophancy in Language Models
- Lang et al. (2024) When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
- Laidlaw et al. (2025) AssistanceZero: Scalably Solving Assistance Games
- Emmons et al. (2025) Observation Interference in Partially Observable Assistance Games
- Ananthakrishnan et al. (2026) Provably Optimal Learning Algorithms for Assistance Games
- Fickinger et al. (2020) Multi-Principal Assistance Games
- Kleine Buening et al. (2025) Strategyproof Reinforcement Learning from Human Feedback
36.3 Preference elicitation in AI and decision analysis #
Preference elicitation is the older name, in AI and decision analysis, for the problem PBO solves: asking a person a few questions in order to make a good decision for them. Its classical results still stand; we found no later work that contradicts them. An efficient allocation of many items can require communication exponential in their number (Nisan and Segal, 2006); whether a class of utilities can be elicited with polynomially many queries depends on its structure (Blum et al., 2004); and elicitation cast as a partially observable Markov decision process, a planning problem over beliefs, has continuous states and actions that put it beyond standard solution techniques (Boutilier, 2002). Choosing the question with the highest expected value of information, the expected improvement in the final decision from hearing the answer, goes back to Chajewska et al. (2000) and is the counterpart of the decision-theoretic acquisition functions of Section 19.4 (inference); minimizing the worst-case regret over utilities consistent with the answers is the main alternative (Wang and Boutilier, 2003).
The link with PBO is exact in one place: introducing qEUBO, Astudillo, Lin, Bakshy, and Frazier wrote that it is closely related to Viappiani and Boutilier (2010), who connected optimal recommendation sets with one-step Bayes-optimal query sets; they also noted that the two criteria choose different queries when answers are noisy, and their appendix uses a lemma deduced from the proof of Theorem 3 in Viappiani and Boutilier's supplementary material (Astudillo et al., 2023) (a journal version appeared later (Viappiani and Boutilier, 2020)). How much can a comparison save? If the learner may ask which of two examples is farther from the decision boundary, under assumptions such as a large margin "it is possible to reveal all the labels of a sample of size using approximately queries", "an exponential improvement over classical active learning", while without them queries are needed in the worst case (Kane et al., 2017) (FOCS 2017), extended to bounded (Massart) noise by Hopkins et al. (2020) (COLT 2020). For CP-nets, a qualitative representation of conditional preferences ("if the main course is fish, I prefer white wine"), Alanazi et al. (2020) (Artificial Intelligence 2020) computed the VC and teaching dimensions (how many examples a learner needs) for learning acyclic CP-nets from examples that differ in one attribute, with near-optimal algorithms even when the oracle errs.
36.3.1 Scale, robustness, and validation with people #
The work closest to PBO is Zintgraf et al. (2018) (AAMAS 2018), who used Gaussian process preference elicitation to choose among the solutions of a multi-objective problem. With simulated and real users, they found that ranking and clustering strategies "outperform the currently used pairwise methods", that "users prefer ranking most", and that monotonicity, through "a linear prior mean at the start and virtual comparisons to the nadir and ideal points" (the worst and best values of every objective), "increases performance"; they demonstrated the framework "in a real-world study on traffic regulation, conducted with the city of Amsterdam".
Elicitation also grew in scale and robustness. Vendrov et al. (2020) (AAAI 2020) wrote the expected value of information as "a differentiable network that can be optimized using gradient methods", and Martin et al. (2024) (IJCAI 2024) learn the response and utility models from data and plan non-myopic elicitation with Monte Carlo tree search. Vayanos et al. (2020), a 2020 preprint whose journal version we did not find, represent uncertainty about the utility as a set and choose queries for worst-case utility or regret; Johnston et al. (2023) (EAAMO 2023) deployed this to prioritize COVID-19 patients for scarce hospital resources with "193 Amazon Mechanical Turk (MTurk) workers", beating random queries by 21% in the utility of the recommended policies. Herin et al. (2024) (ADT 2024) proposed noise-tolerant active learning of the weights of decision aggregation functions, and McElfresh et al. (2021) (AAAI 2021) formalized models of indecision and tested them on a survey about allocating organs. Defresne et al. (2025) (IJCAI 2025) elicit preferences for multi-objective combinatorial problems with "Maximum Likelihood Estimation of a Bradley-Terry preference model" and "an ensemble-based acquisition function inspired from Active Learning", and Bonilla et al. (2026) (ICML 2026) elicit a causal graph from experts with "a three-way likelihood over edge existence and direction" and expected information gain.
36.3.2 What it means for PBO #
Decision-centric query selection, choosing questions for their effect on the final decision rather than for what they reveal about the utility everywhere, appeared in three communities in turn (inference): decision analysis, with Chajewska et al. (2000) and Viappiani and Boutilier (2010); preference-based reinforcement learning, with information-directed reward learning (2021) and query-policy misalignment (2024); and PBO, with qEUBO (2023) acknowledging its debt to Viappiani and Boutilier. We did not check whether PBO papers cite the reinforcement learning work, or the reverse.
Zintgraf et al. had already solved two problems within a Gaussian process preference framework, monotonicity (a linear prior mean plus virtual comparisons with the nadir and ideal points) and query format (ranking beats pairs, and users prefer it), which later PBO work still treats as open (inference): a monotone neural ensemble for preference exploration (Wang et al., 2025a) (NeurIPS 2025) attributes its gain to monotonicity in an ablation, and multi-choice interfaces such as GimmBO and MultiBO (Section 20.1) adopt lists; we did not check whether these papers cite Zintgraf et al. The flow also runs the other way: Defresne et al. adopt the Bradley-Terry likelihood and ensemble acquisition familiar from PBO, without a Gaussian process (inference).
Two practices are worth borrowing (inference). Minimax-regret elicitation gives the worst-case regret over every utility consistent with the answers, though it has traditionally assumed noise-free answers, and Herin et al. show it converging with probabilistic methods; a PBO variant could report the worst-case regret over a Gaussian process credible set next to the expected regret, which suits the high-stakes allocations Vayanos et al. had in mind. And Johnston et al. ran a randomized study with 193 users against random queries, whereas PBO user studies usually have a dozen to a few dozen participants (Chapter 32) and rarely a group answering random queries. Kane et al.'s exponential gain needs a margin or a bounded description, the classical point that structure decides whether elicitation is feasible; in Gaussian process PBO the smoothness of the kernel plays that role (inference).
Sources cited in Section 36.3 21
- Nisan and Segal (2006) The communication requirements of efficient allocations and supporting prices
- Blum et al. (2004) Preference Elicitation and Query Learning
- Boutilier (2002) A POMDP Formulation of Preference Elicitation Problems
- Chajewska et al. (2000) Making Rational Decisions using Adaptive Utility Elicitation
- Wang and Boutilier (2003) Incremental Utility Elicitation with the Minimax Regret Decision Criterion
- Viappiani and Boutilier (2010) Optimal Bayesian Recommendation Sets and Myopically Optimal Choice Query Sets
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Viappiani and Boutilier (2020) On the equivalence of optimal recommendation sets and myopically optimal query sets
- Kane et al. (2017) Active classification with comparison queries
- Hopkins et al. (2020) Noise-tolerant, Reliable Active Classification with Comparison Queries
- Alanazi et al. (2020) The complexity of exact learning of acyclic conditional preference networks from swap examples
- Zintgraf et al. (2018) Ordered Preference Elicitation Strategies for Supporting Multi-Objective Decision Making
- Vendrov et al. (2020) Gradient-based Optimization for Bayesian Preference Elicitation
- Martin et al. (2024) Model-Free Preference Elicitation
- Vayanos et al. (2020) Robust Active Preference Elicitation
- Johnston et al. (2023) Deploying a Robust Active Preference Elicitation Algorithm on MTurk: Experiment Design, Interface, and Evaluation for COVID-19 Patient Prioritization
- Herin et al. (2024) Noise-Tolerant Active Preference Learning for Multicriteria Choice Problems
- McElfresh et al. (2021) Indecision Modeling
- Defresne et al. (2025) Preference Elicitation for Multi-objective Combinatorial Optimization with Active Learning and Maximum Likelihood Estimation
- Bonilla et al. (2026) Causal Preference Elicitation
- Wang et al. (2025a) Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble
36.4 Recommender systems and feedback loops #
A recommender system chooses what a person sees, and what the person then clicks becomes its training data: a feedback loop. A common claim is that such loops trap people in a filter bubble, a narrowing diet of content that reinforces what they already like, and by extension that PBO, which also chooses what a person sees, could form one within a session. The evidence has split. In simulation and theory, training on data confounded by earlier recommendations "homogenizes user behavior without increasing utility" (Chaney et al., 2018) (RecSys 2018); feedback loops amplify popularity bias, and "the impact of feedback loop is generally stronger for the users who belong to the minority group" (Mansoury et al., 2020) (CIKM 2020); a model of users who drift toward what a matrix-factorization recommender shows them yields preference amplification under conditions the authors characterize and validate with simulations (Kalimeris et al., 2021) (KDD 2021); and when preferences move toward what people consume and like, "standard user reward maximization is an almost trivial goal" ("a large class of simple algorithms will achieve only constant regret") (Dean and Morgenstern, 2022) (EC 2022).
Large field experiments find limited short-term effects on polarization. On Facebook, individual choices limited exposure to diverse content more than ranking did (Bakshy et al., 2015); on YouTube, "relying exclusively on the YouTube recommender results in less partisan consumption", and when partisan users switch to moderate content, the sidebar recommender "'forgets' their partisan preference within roughly 30 videos" (Hosseinmardi et al., 2024) (PNAS 2024); and in "four experiments with nearly 9,000 participants", manipulating recommendations to create filter bubbles and rabbit holes "has limited effects on opinions" (Liu et al., 2025b) (PNAS 2025). But in a 7-week randomized experiment with 4,965 users of X in the United States, switching from a chronological feed to the algorithmic one raised the probability that users considered the investigations into Donald Trump unacceptable by 5.5 percentage points and moved policy priorities by 0.11 standard deviations, while switching back had no comparable effect, and neither direction significantly changed affective polarization or partisan identity (Gauthier et al., 2026) (Nature 2026).
Recommenders also face elicitation directly at cold start, when a new user has no history: asking only 2 questions improved recommendations by 25% over a static model, with significant benefits from offline embeddings learned from other users and from bandit-style exploration (Christakopoulou et al., 2016) (KDD 2016), and surveys of conversational recommendation list question-based elicitation and evaluation with simulated users among the open challenges (Jannach et al., 2021; Gao et al., 2021). Critiquing recommenders take directional feedback on attributes ("cheaper", "more like this one but quieter"); Antognini and Faltings (2021) (RecSys 2021) process critiques up to 25.6 times faster than the best baselines with a variational autoencoder. PEBOL is discussed in Section 35.2.1.
36.4.1 What it means for PBO #
All of the following is inference. Warm starts from population embeddings have been standard in recommender cold start since at least 2016, so the meta-learned and population priors of PBO (PABBO, and the prior Liao et al. pretrain from user models; Section 32.5) are the same idea, and that Liao et al.'s prior helped significantly only in early iterations is consistent with this reading. Dean and Morgenstern's result transfers to PBO with changing preferences: if the person's utility moves toward what they are shown, low regret can be achieved trivially, so stationarity or a safe-shift criterion is the meaningful target (Figure 42.1). The field experiments make an in-session filter bubble plausible but suggest that its effect in a short session may be small, and Gauthier et al.'s asymmetry suggests that changes induced by an optimizer might not reverse when it stops; they concern weeks of news feeds, not a design session, so how far they transfer is uncertain. And critiquing is the closest relative of the projective and line-search queries of PBO (Section 20.2, Section 20.3): directional critique is a type of observation PBO could adopt.
Sources cited in Section 36.4 12
- Chaney et al. (2018) How algorithmic confounding in recommendation systems increases homogeneity and decreases utility
- Mansoury et al. (2020) Feedback Loop and Bias Amplification in Recommender Systems
- Kalimeris et al. (2021) Preference Amplification in Recommender Systems
- Dean and Morgenstern (2022) Preference Dynamics Under Personalized Recommendations
- Bakshy et al. (2015) Exposure to ideologically diverse news and opinion on Facebook
- Hosseinmardi et al. (2024) Causally estimating the effect of YouTube’s recommender system using counterfactual bots
- Liu et al. (2025b) Short-term exposure to filter-bubble recommendation systems has limited polarization effects: Naturalistic experiments on YouTube
- Gauthier et al. (2026) The political effects of X’s feed algorithm
- Christakopoulou et al. (2016) Towards Conversational Recommender Systems
- Jannach et al. (2021) A Survey on Conversational Recommender Systems
- Gao et al. (2021) Advances and Challenges in Conversational Recommender Systems: A Survey
- Antognini and Faltings (2021) Fast Multi-Step Critiquing for VAE-based Recommender Systems
36.5 Interactive evolutionary computation #
Interactive evolutionary computation is evolutionary search in which a person, not a formula, judges the fitness of each candidate: the algorithm breeds variants of the designs the person liked, shows them, and repeats. The first comprehensive survey since Takagi (2001) named user fatigue the main challenge and does not compare the field with Bayesian optimization in its abstract (Wang and Pei, 2024) (Applied Soft Computing 2024). Interactive differential evolution based on paired comparisons predates the 2017 formulation of PBO (Takagi and Pallez, 2009) (NaBIC 2009); the latent vector of a generative adversarial network "can be put under evolutionary control" to evolve images toward a target (Bontrager et al., 2018) (EvoMUSART 2018); and quality diversity through human feedback "progressively infers diversity metrics from human judgments of similarity among solutions" and "is more favorably received in user studies" for text-to-image generation (Ding et al., 2024) (ICML 2024). Quality-diversity search returns an archive of good solutions that differ from one another, rather than one best solution. The exoskeleton study of Lee et al. (2023), in which an evolutionary algorithm proposed candidates for a neural ranker and the wearer's forced choices, is in Section 33.1.3.
Evolution strategies have also been compared with Bayesian optimization where a measured cost such as metabolic rate is minimized. Adaptive-sampling CMA-ES, which spends evaluation time where candidates are hard to sort, "converged more efficiently and reliably in complex landscapes, while in simpler landscapes, AS-CMA was less efficient but equally reliable" (Martin and Collins, 2026) (Evolutionary Computation 2026). In simulated exoskeleton tuning on a fitted metabolic landscape (Kutulakos and Slade, 2024) (a 2024 bioRxiv preprint), Bayesian optimization converged near the optimum after about 60 evaluations on a fixed landscape; when the landscape changed as a simulated novice adapted, CMA-ES reached the optimum at similar rates for expert and novice simulations; and the authors conclude that no algorithm is clearly superior for every use case. Both studies use a measured cost, not preferences.
36.5.1 What it means for PBO #
All of the following is inference. Interactive evolution was using surrogate fitness models, pairwise comparisons, and small displays against fatigue before 2017. PBO papers often claim sample efficiency over interactive evolution, but the fair comparison is with surrogate-assisted interactive evolution under human preference feedback, one of the missing studies listed at the end of this chapter. An archive of diverse, high-quality options fits the goal of exploration better than one maximizer when preferences are still forming, and gallery and batch queries in PBO do this only in part. And CMA-ES's robustness to a user who adapts is the scalar counterpart of PBO's drift problem: a population-based or windowed method may tolerate drift better than a Gaussian process posterior that keeps every stale comparison (Section 29.10). Beyond diversity and drift, the advances of interactive evolution itself (new operators, models of fatigue) supply nothing PBO lacks.
Sources cited in Section 36.5 8
- Takagi (2001) Interactive evolutionary computation: fusion of the capabilities of EC optimization and human evaluation
- Wang and Pei (2024) A comprehensive survey on interactive evolutionary computation in the first two decades of the 21st century
- Takagi and Pallez (2009) Paired Comparisons-based Interactive Differential Evolution
- Bontrager et al. (2018) Deep Interactive Evolution
- Ding et al. (2024) Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization
- Lee et al. (2023) User preference optimization for control of ankle exoskeletons using sample efficient active learning
- Martin and Collins (2026) Improving CMA-ES Convergence Speed, Efficiency, and Reliability in Noisy Robot Optimization Problems
- Kutulakos and Slade (2024) Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance
36.6 Learning to rank #
Learning to rank trains models that order items, such as search results, from clicks or judgments; label ranking predicts, for each input, an ordering of a fixed set of labels; and preference learning is the umbrella term. These fields have studied pairwise comparisons for decades, and three of their results bear on how PBO should collect and read its data.
36.6.1 How much a comparison is worth #
Two claims about comparisons circulate: that a comparison carries at most one bit per query, far less than a rating (Section 16.6), and that pairwise feedback is not less efficient in an information-theoretic sense. Both can be true. Shah et al. (2016) (JMLR 2016) derived tight minimax bounds (bounds on the best achievable worst-case error) for estimating item qualities under the Bradley-Terry-Luce and Thurstone models, which "depend on the topology of the comparison graph induced by the subset of pairs being compared, via the spectrum of the Laplacian of the comparison graph", and "the error rates in the ordinal and cardinal settings have identical scalings apart from constant pre-factors". Each comparison carries less information, but the rate is of the same order.
36.6.2 Three results PBO rarely uses #
Parametric links buy little. Heckel et al. (2019) (Annals of Statistics 2019) analyzed an algorithm that counts the comparisons each item wins and chooses the next pair by confidence intervals, proved that it recovers the ranking "using a number of comparisons that is optimal up to logarithmic factors" without any parametric model, and settled what they call "a long-standing open question": parametric assumptions such as the Thurstone or Bradley-Terry-Luce models yield at most logarithmic gains for stochastic comparisons.
Fitted utilities are not always majority-respecting. Noothigattu et al. (2020) (NeurIPS 2020) showed that "a large class of random utility models (including the Thurstone-Mosteller Model), when estimated using the MLE, satisfy a Pareto efficiency condition" and "a strong monotonicity property", but "fail certain other consistency conditions from social choice theory, and in particular do not always follow the majority opinion".
Position bias can be measured and corrected. Joachims et al. (2017) (WSDM 2017) started from the observation that "position bias in search rankings strongly influences how many clicks a result receives". Their counterfactual framework treats the probability that a result in a given position is examined as a propensity, reweights clicks by its inverse to obtain unbiased learning to rank, and is "robust to noise and propensity model misspecification". The propensity can be estimated only if the position of results is varied, for instance by randomizing it for some users.
On label ranking, Thies et al. (2026a) (ICML 2026) proved that full-ranking calibration implies the other notions, that "sub-ranking and top-k calibration are incomparable", and that "popular label ranking models are often poorly calibrated"; a model is calibrated when events it predicts with probability happen about a fraction of the time. MORE-PLR (Thies et al., 2026b) (Machine Learning 2026) predicts partial label rankings, in which tied labels share a bucket.
36.6.3 Position bias in a preference loop #
The counterpart in PBO of a ranking position is the left or right slot of a pair, or the position in a gallery. In the human studies of PBO that we read, we saw none that reported randomizing or modeling the order of presentation, though we did not check the methods section of every paper; language-model judges have strong position bias, and averaging over both orders is the simplest remedy (Section 35.2.4) (Wang et al., 2024a). Figure 36.1 shows what an uncorrected position bias can do to a preference loop.
Things to try:
- Leave Incumbent always left selected with and press play. The answers favor the incumbent about three times in four even when the challenger is better, the model reads the slot's advantage as quality, and the magenta curve stalls well above zero. The line "Answers for the left slot" mixes quality and slot advantage, and no model fitted to these data can separate them: the confounding Joachims et al. break by varying the position.
- Switch to Random side. The bias now favors incumbent and challenger equally often, so it acts as extra noise rather than a thumb on the scale, and the green curve keeps falling.
- Switch to Random side, bias modeled. The likelihood gains one parameter, the slot preference, whose estimate appears next to the true ; in these runs modeling mainly buys a measurement, with lower regret than random sides only at larger .
- Drag to 0. All three interfaces behave about the same: the same-slot design is harmless only if the person has no slot preference, and the same-slot data cannot tell you whether they do.
A bias toward the challenger's slot (not shown) did little harm in these simulations, because it makes the loop switch incumbents too eagerly, which costs little when the optimizer keeps the best compared design; the asymmetry is a property of this incumbent-and-challenger loop, not a general result (inference). The practical rule follows (inference): randomize the side of every pair, and record it, so that a slot preference can be measured and modeled later.
36.6.4 What it means for PBO #
All of the following is inference. By Heckel et al., the sample efficiency of PBO should come mainly from the kernel sharing information between nearby inputs, not from the form of the link. By Shah et al., the error depends on the spectrum of the comparison graph, and a rule that always compares against the current best builds a star-shaped graph that concentrates information on the incumbent: good for locating the maximum, but not necessarily for learning a reusable utility. Two 2026 papers reached similar conclusions inside PBO (Section 27.6, Figure 27.1): Shao et al. (2026), a preprint, found that EUBO selects pairs that form isolated components of the comparison graph, making the Laplace likelihood Hessian rank deficient, and Pukdee et al. (2026) (ICML 2026) identified the margin and the connectivity of the comparison graph as what governs the sample efficiency of Bradley-Terry learning; this matches the distinction of Section 35.3.6 between pairs that locate the optimum and pairs that teach a utility. PBO papers rarely report whether their predicted pairwise probabilities are calibrated on held-out human comparisons, although Gaussian approximations predict duel outcomes poorly (Takeno et al., 2023) (Section 27.4), and label ranking offers ready-made metrics. And a Gaussian process utility pooled across users is a maximum likelihood random utility model, so by Noothigattu et al. it may contradict the majority on some pairs, which adds to the Borda-count result of Section 35.4.4.
Sources cited in Section 36.6 10
- Shah et al. (2016) Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
- Heckel et al. (2019) Active ranking from pairwise comparisons and when parametric assumptions do not help
- Noothigattu et al. (2020) Axioms for Learning from Pairwise Comparisons
- Joachims et al. (2017) Unbiased Learning-to-Rank with Biased Feedback
- Thies et al. (2026a) Calibrated Preference Learning: The Case of Label Ranking
- Thies et al. (2026b) MORE-PLR: multi-output regression employed for partial label ranking
- Wang et al. (2024a) Large Language Models are not Fair Evaluators
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Pukdee et al. (2026) What Does Preference Learning Recover from Pairwise Comparison Data?
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
36.7 Automated science and self-driving labs #
A self-driving lab couples robotic experimentation with an optimizer, usually Bayesian optimization, that chooses the next experiment. One laboratory study uses a person's judgment as the only measurement: Deneault et al. (2025) (Digital Discovery 2025) tuned a 3-D printer for a printing goal that is "difficult to measure with sensors but can be readily evaluated from human judgment"; we could read only the abstract, so we do not report its campaign sizes, query format, or gain. Other laboratory applications with experts in the loop are in Section 34.2.
Adesiji et al. (2026) (Digital Discovery 2026) define the acceleration factor, the ratio of the experiments a reference strategy needs to reach a target to those the optimizer needs, and the enhancement factor, the gain after a fixed number of experiments. Across 42 studies and 63 benchmarks the median reported acceleration factor was 6 (range 1.3 to 100), and the enhancement factor peaked at about 10 to 20 experiments per dimension. As they report, Shields et al. (2021) (Nature 2021) found that by the 15th experiment Bayesian optimization's average performance exceeded that of 50 expert chemists (Adesiji et al. call the reaction space ten-dimensional; the published data set has five choices and 1,728 measured conditions, Section 23.1); Chapter 23 replays such an optimization, and Section 23.4 returns to the chemists.
People in automated experiments mostly supervise. "The likely strategy for the next several years will be human-in-the-loop automated experiments", in which a "human operator monitors experiment progression" and adjusts the agent's policy (Kalinin et al., 2024) (a 2023 preprint later published in Microscopy Today); with deep kernel learning, "for certain parameter combinations the experiment path can be trapped in the local minima", and monitoring was used to construct "intervention strategies" (Pratiush et al., 2025) (a 2024 preprint later published in Digital Discovery); human input to an autonomous synthesis agent improved sampling efficiency on synthetic benchmarks and found processing regions that stabilize metastable phases in real Bi-Ti-O thin-film experiments (Chang et al., 2026) (SARA-H, a 2026 preprint); and the GIFTERS checklist for trustworthy AI in materials discovery (median score 5 out of 7 across the reviewed work) argues for keeping people in the loop (Amirian et al., 2025) (a 2025 preprint). LGBO used a language model's preferences about regions as side information for scalar Bayesian optimization in a wet-lab experiment (Section 35.2.2).
36.7.1 What it means for PBO #
All of the following is inference. People in self-driving labs mostly act as supervisors, stepping in when the surrogate gets stuck or the objective turns out to be wrong, and pairwise preference is one channel among several, closer to the BO-as-assistant arrangement of Section 32.2 than to a loop in which the person only compares. The acceleration and enhancement factors, and the peak at 10 to 20 experiments per dimension, offer PBO a reporting standard: PBO papers mostly report regret on synthetic functions and rarely an acceleration factor with real users against a human-only or random baseline, Deneault et al. being a rare exception. Expert pairwise input pays off when the expert's judgment carries information the surrogate lacks, such as a subjective quality (Deneault et al.) or an unmeasured property, as with the materials experts of Mikkola et al. (2020) (Section 34.2.1), and not when the objective can be measured directly, where Bayesian optimization beat 50 expert chemists on average (Shields et al., 2021).
Sources cited in Section 36.7 8
- Deneault et al. (2025) Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities
- Adesiji et al. (2026) Benchmarking self-driving labs
- Shields et al. (2021) Bayesian reaction optimization as a tool for chemical synthesis
- Kalinin et al. (2024) Human-in-the-loop: The future of Machine Learning in Automated Electron Microscopy
- Pratiush et al. (2025) Building Workflows for Interactive Human in the Loop Automated Experiment (hAE) in STEM-EELS
- Chang et al. (2026) Autonomous Materials Exploration by Integrating Automated Phase Identification and AI-Assisted Human Reasoning
- Amirian et al. (2025) Building Trustworthy AI for Materials Discovery: From Autonomous Laboratories to Z-scores
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
36.8 Common claims, checked #
Table 36.1 checks claims about these neighboring fields that are repeated in writing about PBO; two more, on reward-model calibration and on how many comparisons each field uses, are checked in Section 35.5.
| Claim | What the evidence says |
|---|---|
| Inverse reinforcement learning and AI alignment treat the reward as fixed but unknown, so AI independently adopted PBO's assumption of a stable latent utility. | Accurate as a description of the formalisms, but questioned inside the field: "existing AI alignment approaches assume that preferences are static, which is unrealistic", which "may undermine the soundness of existing alignment techniques" (Carroll et al., 2024), and the usual model of how people generate preferences from a fixed reward is flawed (Knox et al., 2024). The shared assumption is a modeling convention, not independent evidence that stable utilities exist. |
| Potential-based reward shaping keeps instrumental queries from distorting inference about terminal preferences. | Shaping is one of the transformations that behavioral data cannot detect: by Theorem 3.3 of Skalse et al. (2023) (ICML 2023), Boltzmann-rational policies determine the reward only up to S′-redistribution and potential shaping. We found no source that uses shaping to design preference queries; the claim is an analogy. |
| HERON and DIPPER are published methods for hierarchical preference design in reinforcement learning. | Correct: HERON appeared at ICML 2025 and DIPPER, now titled "Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach", at ICLR 2026 (Bukharin et al., 2023; Singh et al., 2024). |
| Hejna et al. (CoRL 2023) used meta-learned reward functions to need 20 times fewer queries than PEBBLE. | The paper is Hejna and Sadigh (2022), in PMLR volume 205 (the 6th Conference on Robot Learning, December 14 to 18, 2022, published March 6, 2023). Its abstract reports reducing online feedback in Meta-World "by 20×", with a real Franka Panda demonstration, but does not name PEBBLE as the comparison. |
| Cooperative inverse reinforcement learning treats the human's reward parameters as existing, stable, and unknown to the robot. | Accurate. AssistanceZero keeps this assumption (Laidlaw et al., 2025); Emmons et al. (2025) add partial observability. |
| Sequential strategic misreporting remains possible, so acquisition should be incentive-compatible. | Now quantified for RLHF: one strategic labeler can cause arbitrarily large misalignment, and any strategyproof algorithm can be times worse than optimal (Kleine Buening et al., 2025). |
| Exact elicitation, cast as a POMDP, is PSPACE-hard (attributed to Boutilier 2002), and the classical complexity results set the limits of elicitation. | The classical results stand (Section 36.3), and no newer CP-net work relevant to PBO turned up in our searches. The PSPACE statement is not in Boutilier (2002), whose abstract says standard POMDP techniques cannot solve the problem because its states and actions are continuous; the PSPACE-completeness of finite-horizon partially observed Markov decision problems with finitely many states (Papadimitriou and Tsitsiklis, 1987) is not a result about elicitation. |
| Interactive evolution makes discovery possible without domain knowledge, and design galleries implicitly acknowledged that preferences are discovered. | Holds. Recent interactive quality-diversity work states the same goals, few alternatives "to reduce cognitive load" that "should be diverse but similar to the previous user selection, to reduce user fatigue", and its windowed MAP-Elites (which keeps the best solution in each cell of a grid of behaviors) "finds more appropriate solutions to the user's taste", tested with "controllable artificial users" (Sfikas et al., 2023). |
Sources cited in Section 36.8 12
- Carroll et al. (2024) AI Alignment with Changing and Influenceable Reward Functions
- Knox et al. (2024) Models of human preference for learning reward functions
- Skalse et al. (2023) Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
- Bukharin et al. (2023) Deep Reinforcement Learning from Hierarchical Preference Design
- Singh et al. (2024) Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
- Hejna and Sadigh (2022) Few-Shot Preference Learning for Human-in-the-Loop RL
- Laidlaw et al. (2025) AssistanceZero: Scalably Solving Assistance Games
- Emmons et al. (2025) Observation Interference in Partially Observable Assistance Games
- Kleine Buening et al. (2025) Strategyproof Reinforcement Learning from Human Feedback
- Boutilier (2002) A POMDP Formulation of Preference Elicitation Problems
- Papadimitriou and Tsitsiklis (1987) The Complexity of Markov Decision Processes
- Sfikas et al. (2023) Controllable Exploration of a Design Space via Interactive Quality Diversity
36.9 What the neighbors solved first #
Table 36.2 lists problems that a neighboring field has solved or given a ready practice for, and that the PBO literature still treats as open or has not adopted, with whether the suggestion depends on the Gaussian process framework.
| Problem | Field that got there first | Status in PBO | Depends on the GP framework? |
|---|---|---|---|
| Decision-centric query selection | decision analysis (Viappiani and Boutilier 2010); preference-based RL (Lindner et al. 2021, Hu et al. 2024) | adopted by qEUBO (2023), source acknowledged | no |
| Monotonicity and ranking queries | GP preference elicitation (Zintgraf et al. 2018) | still treated as open | monotonicity by virtual comparisons: yes; ranking: no |
| Simulated teachers with realistic irrationalities | preference-based RL (B-Pref 2021) | benchmarks mostly use homoscedastic, independent noise | no |
| Measuring and correcting position bias | learning to rank (Joachims et al. 2017); language-model judges (Wang et al. 2024) | no reports of randomized or modeled order found | no |
| Warm starts from population embeddings | recommender systems (Christakopoulou et al. 2016) | population priors presented as new | no |
| Designing the comparison graph | ranking estimation (Shah et al. 2016) | raised in 2026, through the rank-deficient Laplace Hessian | partly |
| Calibration metrics | label ranking (Thies et al. 2026) | calibration of pairwise probabilities rarely reported | no |
| Worst-case guarantees | robust elicitation (Vayanos et al. 2020, Johnston et al. 2023, Herin et al. 2024) | probabilistic posterior only | yes, over a GP credible set |
| Robustness when the user adapts | evolution strategies (Martin and Collins 2026, Kutulakos and Slade 2024) | no drifting-utility model | no |
| An archive of diverse options | quality diversity through human feedback (Ding et al. 2024) | no learned diversity objective | no |
| Directional feedback on attributes | critiquing recommenders (for example, Antognini et al. 2021) | not used as an observation | no |
| Formalizing the system's influence on the person | AI safety (Carroll et al. 2022, 2024; Emmons et al. 2025) | not measured | no |
| Reporting by acceleration factors | self-driving labs (Adesiji et al. 2026) | mostly regret on synthetic functions | no |
Most of these suggestions come from the problem itself, not from the Gaussian process framework; only monotonicity through virtual comparisons, worst-case regret over a credible set, and the rank deficiency of the comparison graph are repairs inside it.
36.10 Settled, contested, missing #
Settled. Decision-centric query selection was proposed in decision analysis, then in preference-based reinforcement learning, then in PBO, and qEUBO acknowledges its decision-analytic source (Astudillo et al., 2023). Pairwise feedback carries less information per query than ratings but achieves the same rate up to constants (Shah et al., 2016), and parametric links add at most logarithmic gains for ranking (Heckel et al., 2019). The choice of model for how people generate preferences changes what can be learned (Knox et al., 2024), and small adversarial biases can produce arbitrarily large errors in inferred rewards (Hong et al., 2023). Position bias in rankings is measurable and correctable when position is varied (Joachims et al., 2017). Simulated teachers with perfect preferences are unrealistic (Lee et al., 2021a). Systems trained with long horizons have an incentive to shift preferences (Carroll et al., 2022), and manipulation of a vulnerable minority has been demonstrated in language models (Williams et al., 2025).
Contested. Whether recommender feedback loops change attitudes: simulations say yes, most field experiments find limited short-term effects, and one 7-week experiment found an asymmetric effect (Gauthier et al., 2026). Whether evolution strategies or Bayesian optimization suit human-in-the-loop optimization better: the answer depends on the landscape and on whether the user adapts (Martin and Collins, 2026; Kutulakos and Slade, 2024). Whether expert pairwise input improves on a well-specified optimizer: it helped when the goal was subjective (Deneault et al., 2025) and lost to the optimizer when the yield could be measured (Shields et al., 2021).
Missing. A measurement of preference change caused by an acquisition function in a study of PBO with people. Any study of PBO that randomizes or models presentation order. A comparison of interactive evolution and PBO under pairwise preference feedback, including one with surrogate-assisted interactive evolution as the baseline, and a quantitative model of fatigue taken from interactive evolution. A PBO method that optimizes a diversity objective learned from people, uses directional critique as an observation, or models a drifting utility. Worst-case guarantees reported alongside the posterior. Acceleration factors measured with real users against human-only or random baselines. A self-driving lab with more than one human oracle giving pairwise preferences in a closed loop. Benchmarks that use B-Pref-style irrational simulated users.
Sources cited in Section 36.10 14
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Shah et al. (2016) Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
- Heckel et al. (2019) Active ranking from pairwise comparisons and when parametric assumptions do not help
- Knox et al. (2024) Models of human preference for learning reward functions
- Hong et al. (2023) On the Sensitivity of Reward Inference to Misspecified Human Models
- Joachims et al. (2017) Unbiased Learning-to-Rank with Biased Feedback
- Lee et al. (2021a) B-Pref: Benchmarking Preference-Based Reinforcement Learning
- Carroll et al. (2022) Estimating and Penalizing Induced Preference Shifts in Recommender Systems
- Williams et al. (2025) On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
- Gauthier et al. (2026) The political effects of X’s feed algorithm
- Martin and Collins (2026) Improving CMA-ES Convergence Speed, Efficiency, and Reliability in Noisy Robot Optimization Problems
- Kutulakos and Slade (2024) Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance
- Deneault et al. (2025) Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities
- Shields et al. (2021) Bayesian reaction optimization as a tool for chemical synthesis
Further reading #
- Lee et al. (2021a), the B-Pref benchmark, is the clearest catalog of the ways a simulated person can be irrational, and reads as a checklist for PBO benchmarks.
- Knox et al. (2024) and Hong et al. (2023) together explain why the model of how a person answers matters as much as the model of what they want.
- Carroll et al. (2024) is the most direct treatment of systems that change the preferences they learn from.
- Zintgraf et al. (2018) is a Gaussian process preference elicitation study with real users that settled monotonicity and query format in 2018.
- Joachims et al. (2017) shows how a biased presentation can be turned into a measured propensity, the idea behind Figure 36.1.
- Adesiji et al. (2026) defines the acceleration and enhancement factors and collects them across self-driving lab studies.
References
- (2026). Benchmarking self-driving labs. Digital Discovery. Cited in §36.7
- (2020). The complexity of exact learning of acyclic conditional preference networks from swap examples. Artificial Intelligence. doi:10.1016/j.artint.2019.103182. Cited in §36.3
- (2025). Building Trustworthy AI for Materials Discovery: From Autonomous Laboratories to Z-scores. arXiv. preprint Cited in §36.7
- (2026). Provably Optimal Learning Algorithms for Assistance Games. arXiv (a 2026 AI4GOOD Workshop version also exists). preprint Cited in §36.2
- (2021). Fast Multi-Step Critiquing for VAE-based Recommender Systems. Fifteenth ACM Conference on Recommender Systems. doi:10.1145/3460231.3474249. Cited in §36.4
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §36.3 §36.10
- (2015). Exposure to ideologically diverse news and opinion on Facebook. Science. Cited in §36.4
- (2004). Preference Elicitation and Query Learning. Journal of Machine Learning Research. Cited in §36.3
- (2026). Causal Preference Elicitation. ICML 2026 (per OpenReview). Cited in §36.3
- (2018). Deep Interactive Evolution. EvoMUSART 2018. Cited in §36.5
- (2002). A POMDP Formulation of Preference Elicitation Problems. Proceedings of the Eighteenth National Conference on Artificial Intelligence (AAAI-02). Cited in §36.3 §36.8
- (2023). Deep Reinforcement Learning from Hierarchical Preference Design. International Conference on Machine Learning (ICML 2025). Cited in §36.8
- (2022). Estimating and Penalizing Induced Preference Shifts in Recommender Systems. ICML 2022. Cited in §36.2 §36.10
- (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. Cited in §36.2 §36.8
- (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. TMLR 2023. Cited in §36.1
- (2000). Making Rational Decisions using Adaptive Utility Elicitation. Proceedings of the Seventeenth National Conference on Artificial Intelligence (AAAI-00). Cited in §36.3
- (2018). How algorithmic confounding in recommendation systems increases homogeneity and decreases utility. Proceedings of the 12th ACM Conference on Recommender Systems. Cited in §36.4
- (2026). Autonomous Materials Exploration by Integrating Automated Phase Identification and AI-Assisted Human Reasoning. arXiv. preprint Cited in §36.7
- (2024). RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences. ICML 2024. Cited in §36.1
- (2024). Listwise Reward Estimation for Offline Preference-based Reinforcement Learning. ICML 2024. Cited in §36.1
- (2016). Towards Conversational Recommender Systems. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Cited in §36.4
- (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems. Cited in §36.1
- (2022). Preference Dynamics Under Personalized Recommendations. EC 2022. Cited in §36.4
- (2025). Preference Elicitation for Multi-objective Combinatorial Optimization with Active Learning and Maximum Likelihood Estimation. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. Cited in §36.3
- (2025). Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities. Digital Discovery. Cited in §36.7 §36.10
- (2024). Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization. ICML 2024. Cited in §36.5
- (2025). Observation Interference in Partially Observable Assistance Games. ICML 2025. Cited in §36.2 §36.8
- (2023). User Tampering in Reinforcement Learning Recommender Systems. AIES 2023. Cited in §36.2
- (2020). Multi-Principal Assistance Games. arXiv. preprint Cited in §36.2
- (2021). Advances and Challenges in Conversational Recommender Systems: A Survey. AI Open. Cited in §36.4
- (2023). Scaling Laws for Reward Model Overoptimization. Proceedings of the 40th International Conference on Machine Learning (ICML 2023). Cited in §36.1
- (2026). The political effects of X’s feed algorithm. Nature. doi:10.1038/s41586-026-10098-2. Cited in §36.4 §36.10
- (2017). Inverse Reward Design. NeurIPS 2017. Cited in §36.2
- (2025). Influencing Humans to Conform to Preference Models for RLHF. Transactions on Machine Learning Research. Cited in §36.1
- (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. Cited in §36.6 §36.10
- (2022). Few-Shot Preference Learning for Human-in-the-Loop RL. Conference on Robot Learning. Cited in §36.8
- (2023). Inverse Preference Learning: Preference-based RL without a Reward Function. NeurIPS 2023. Cited in §36.1
- (2024). Contrastive Preference Learning: Learning from Human Feedback without RL. ICLR 2024. Cited in §36.1
- (2024). Noise-Tolerant Active Preference Learning for Multicriteria Choice Problems. Algorithmic Decision Theory. Cited in §36.3
- (2023). On the Sensitivity of Reward Inference to Misspecified Human Models. ICLR 2023. Cited in §36.1 §36.10
- (2020). Noise-tolerant, Reliable Active Classification with Comparison Queries. Conference on Learning Theory. Cited in §36.3
- (2024). Causally estimating the effect of YouTube’s recommender system using counterfactual bots. Proceedings of the National Academy of Sciences. Cited in §36.4
- (2024). Query-Policy Misalignment in Preference-Based Reinforcement Learning. ICLR 2024. Cited in §36.1
- (2021). A Survey on Conversational Recommender Systems. ACM Computing Surveys. Cited in §36.4
- (2017). Unbiased Learning-to-Rank with Biased Feedback. WSDM 2017. Cited in §36.6 §36.10
- (2023). Deploying a Robust Active Preference Elicitation Algorithm on MTurk: Experiment Design, Interface, and Evaluation for COVID-19 Patient Prioritization. EAAMO 2023. Cited in §36.3
- (2021). Preference Amplification in Recommender Systems. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. Cited in §36.4
- (2024). Human-in-the-loop: The future of Machine Learning in Automated Electron Microscopy. Microscopy Today. doi:10.1093/mictod/qaad096. Cited in §36.7
- (2017). Active classification with comparison queries. FOCS 2017. Cited in §36.3
- (2024). Goodhart's Law in Reinforcement Learning. ICLR 2024. Cited in §36.2
- (2025). Strategyproof Reinforcement Learning from Human Feedback. NeurIPS 2025. Cited in §36.2 §36.8
- (2024). Models of human preference for learning reward functions. TMLR 2024. Cited in §36.1 §36.8 §36.10
- (2024). Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance. bioRxiv. preprint Cited in §36.5 §36.10
- (2024). Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification. NeurIPS 2024. Cited in §36.2
- (2025). AssistanceZero: Scalably Solving Assistance Games. ICML 2025. Cited in §36.2 §36.8
- (2024). When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback. NeurIPS 2024. Cited in §36.2
- (2021a). B-Pref: Benchmarking Preference-Based Reinforcement Learning. NeurIPS 2021 Datasets and Benchmarks. Cited in §36.1 §36.10
- (2021b). PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training. ICML 2021. Cited in §36.1
- (2023). User preference optimization for control of ankle exoskeletons using sample efficient active learning. Science Robotics. Cited in §36.5
- (2021). Information Directed Reward Learning for Reinforcement Learning. NeurIPS 2021. Cited in §36.1
- (2025b). Short-term exposure to filter-bubble recommendation systems has limited polarization effects: Naturalistic experiments on YouTube. Proceedings of the National Academy of Sciences. Cited in §36.4
- (2020). Feedback Loop and Bias Amplification in Recommender Systems. CIKM 2020. Cited in §36.4
- (2026). Improving CMA-ES Convergence Speed, Efficiency, and Reliability in Noisy Robot Optimization Problems. Evolutionary Computation. Cited in §36.5 §36.10
- (2024). Model-Free Preference Elicitation. Thirty-Third International Joint Conference on Artificial Intelligence. Cited in §36.3
- (2021). Indecision Modeling. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v35i7.16746. Cited in §36.3
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §36.7
- (2006). The communication requirements of efficient allocations and supporting prices. Journal of Economic Theory. Cited in §36.3
- (2020). Axioms for Learning from Pairwise Comparisons. Advances in Neural Information Processing Systems. Cited in §36.6
- (2020). Dueling Posterior Sampling for Preference-Based Reinforcement Learning. Conference on Uncertainty in Artificial Intelligence. Cited in §36.1
- (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. ICLR 2022. Cited in §36.2
- (1987). The Complexity of Markov Decision Processes. Mathematics of Operations Research. Cited in §36.8
- (2025). Building Workflows for Interactive Human in the Loop Automated Experiment (hAE) in STEM-EELS. Digital Discovery. doi:10.1039/d5dd00033e. Cited in §36.7
- (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. Cited in §36.6
- (2024). Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms. NeurIPS 2024. Cited in §36.1
- (2023). Dueling RL: Reinforcement Learning with Trajectory Preferences. International Conference on Artificial Intelligence and Statistics. Cited in §36.1
- (2023). Controllable Exploration of a Design Space via Interactive Quality Diversity. arXiv (parts published at GECCO 2023). preprint Cited in §36.8
- (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Cited in §36.6 §36.10
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §36.6
- (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. Cited in §36.2
- (2021). Bayesian reaction optimization as a tool for chemical synthesis. Nature. Cited in §36.7 §36.10
- (2024). Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach. International Conference on Learning Representations (ICLR 2026). Cited in §36.8
- (2022). Defining and Characterizing Reward Hacking. Advances in Neural Information Processing Systems 35 (NeurIPS 2022). Cited in §36.2
- (2023). Invariance in Policy Optimisation and Partial Identifiability in Reward Learning. ICML 2023. Cited in §36.8
- (2006). The Optimizer’s Curse: Skepticism and Postdecision Surprise in Decision Analysis. Management Science. Cited in §36.2
- (2001). Interactive evolutionary computation: fusion of the capabilities of EC optimization and human evaluation. Proceedings of the IEEE. Cited in §36.5
- (2009). Paired Comparisons-based Interactive Differential Evolution. NaBIC 2009. Cited in §36.5
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §36.6
- (2026a). Calibrated Preference Learning: The Case of Label Ranking. International Conference on Machine Learning (ICML 2026). Cited in §36.6
- (2026b). MORE-PLR: multi-output regression employed for partial label ranking. Machine Learning 115. Cited in §36.6
- (2020). Robust Active Preference Elicitation. arXiv (journal version not found). preprint Cited in §36.3
- (2020). Gradient-based Optimization for Bayesian Preference Elicitation. AAAI 2020. Cited in §36.3
- (2010). Optimal Bayesian Recommendation Sets and Myopically Optimal Choice Query Sets. Advances in Neural Information Processing Systems. Cited in §36.3
- (2020). On the equivalence of optimal recommendation sets and myopically optimal query sets. Artificial Intelligence. Cited in §36.3
- (2003). Incremental Utility Elicitation with the Minimax Regret Decision Criterion. Proceedings of the Eighteenth International Joint Conference on Artificial Intelligence (IJCAI-03). Cited in §36.3
- (2024). A comprehensive survey on interactive evolutionary computation in the first two decades of the 21st century. Applied Soft Computing. doi:10.1016/j.asoc.2024.111950. Cited in §36.5
- (2023a). Is RLHF More Difficult than Standard RL? NeurIPS 2023. Cited in §36.1
- (2024a). Large Language Models are not Fair Evaluators. ACL 2024. Cited in §36.6
- (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. Cited in §36.3
- (2025). Language Models Learn to Mislead Humans via RLHF. ICLR 2025. Cited in §36.2
- (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Cited in §36.2 §36.10
- (2024). Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback. ICLR 2024. Cited in §36.1
- (2020). Consequences of Misaligned AI. NeurIPS 2020. Cited in §36.2
- (2018). Ordered Preference Elicitation Strategies for Supporting Multi-Objective Decision Making. AAMAS 2018. Cited in §36.3