Preferences and Large Language Models
The most visible use of pairwise comparisons in computing today is not preferential Bayesian optimization (PBO). It is the alignment of large language models, where people (and increasingly other models) read two responses to the same prompt and say which is better, millions of times, and the answers are used to tune the model. The question asked of the annotator is the question Chapter 19 asks of a person tuning a design: which of these two do you prefer? The likelihood used to learn from the answer is, in most alignment work, the Bradley-Terry model of Section 16.4.
From 2023 to September 2026 the two fields met in two directions. Language models moved into preference loops, as priors, simulated users, translators from text to preference labels, feature extractors, and generators of candidates. In the other direction, tools from dueling bandits and Bayesian active learning were applied to prompt optimization, to choosing among outputs and models, and to choosing which comparisons to collect for training. For this book's argument, alignment is where the measurement question was asked at scale: how far a model judge departs from a person, and which comparisons are worth asking for (Section 45.1). The reported gains range from 1% to 6% in win rate to roughly an order of magnitude in labels saved, but most results replace people with a reward-model simulator or a language-model judge, and some careful studies find that choosing comparisons cleverly barely beats choosing them at random. The chapter explains how alignment learns from comparisons, follows the two directions, separates the likelihood the fields share from seven analogies that break, and ends with what has flowed back.
35.1 How language models learn from comparisons #
A reader who knows Gaussian process preference learning already knows most of the machinery. This section introduces the vocabulary of alignment in that reader's terms.
35.1.1 Policies, rewards, and the Bradley-Terry reward model #
A language model, given a prompt , produces a response by sampling words one at a time, which defines a probability distribution over responses, . Alignment research borrows the word policy from reinforcement learning for this distribution: a rule for choosing actions, where the action is the entire response. The model before preference tuning, usually already fine-tuned on written demonstrations, is the reference policy .
A reward model is a separate network that maps a prompt and a response to a single number , often a copy of the language model with its output layer replaced by one scalar (Ouyang et al., 2022). It plays exactly the role of the latent utility in Chapter 18: nobody observes it, and it is learned only from comparisons. If an annotator is shown responses and to prompt , the Bradley-Terry model says
where is the logistic function (language-model papers write it , a letter this book reserves for standard deviations). A reward model is trained by maximizing the log-likelihood of comparisons, each a prompt with a preferred response and a rejected one :
Set beside Equation (19.1), the only differences are the link (logistic rather than probit, two noise distributions of the same random utility family, Section 16.5), the absence of a prior on , and the scale: a large neural network fitted to tens of thousands to millions of comparisons instead of a Gaussian process fitted to a few dozen. The ranking extension, in which the probability of an ordering of responses is a product of softmax choices from the remaining items, is the Plackett-Luce model.
35.1.2 RLHF and the KL-regularized objective #
Reinforcement learning from human feedback (RLHF) uses such a reward model to tune a policy. Christiano et al. (2017) introduced it for games and simulated robots, and Ouyang et al. (2022) applied it to instruction-following language models. It samples responses and collects human comparisons, fits a reward model with Equation (35.2), and then fine-tunes the policy to maximize, for each prompt and on average over prompts,
where and are short for and . The second term is the KL divergence of Section 6.2, how far the tuned policy has moved from the reference, and sets how much movement the reward must pay for. Without it, the policy would drift toward whatever text the reward model overrates, a failure called reward over-optimization: the learned reward keeps rising while true quality falls (Coste et al., 2024). Notice what Equation (35.3) asks for: not the single best response, but a distribution over responses that scores well on average while staying close to the reference, the first analogy that breaks in Section 35.4.2.
35.1.3 DPO and the implicit reward #
Direct preference optimization (DPO) removes the separate reward model and the reinforcement learning step. Rafailov et al. (2023) observed that Equation (35.3) has a closed-form optimum, and that the Bradley-Terry likelihood can then be written directly in terms of the policy.
Fix one prompt and drop it from the notation. Write for the policy and sums over all responses .
- The objective of Equation (35.3) for this prompt is , by the definition of the expectation and of the KL divergence.
- Define and the distribution . Factoring out of both terms of step 1 gives . Writing inside the logarithm splits it into , and because the constant contributes once: .
- A KL divergence is never negative and is zero only when the two
distributions are equal, so is maximized by :(35.4)
- Take logarithms of Equation (35.4) and solve for the reward:(35.5)
- Restore the prompt and substitute Equation (35.5) into the Bradley-Terry model Equation (35.1). The difference contains twice with opposite signs, because depends on the prompt but not on the response, so it cancels. The intractable sum over all possible responses disappears.
- What remains is a likelihood in which the policy itself plays the role of
the reward. Write
for this policy-defined reward. Maximizing the Bradley-Terry likelihood over
a data set of comparisons is then the DPO loss:(35.6)
In the words of its title, the policy "is secretly a reward model": the quantity is called the implicit reward, and Equation (35.6) is Equation (35.2) with in place of , with a Plackett-Luce version for rankings in the paper. Two relatives recur below, both members of one family of convex losses (Tang et al., 2024): IPO (Azar et al., 2024) uses a squared loss, and SLiC a hinge loss. And Equation (35.4) says that the optimal policy is the reference policy reweighted by , a tilt of what the model already does rather than a jump to the best response.
35.1.4 Where comparisons come from #
Offline preference learning uses a fixed data set gathered in advance; online learning samples new responses from the current policy during training and sends them for labeling; active learning also chooses which prompts and pairs are worth labeling. That choice is an acquisition function in the sense of Chapter 12, and it is where tools from PBO and dueling bandits enter alignment. A dueling bandit (Section 21.1) chooses two arms per round and observes which one wins; a contextual dueling bandit first observes a context, here the prompt. Much of the evidence below uses an LLM judge, a language model prompted to compare two responses in place of a person, and reports a win rate, the fraction of head-to-head comparisons against a fixed baseline model that the tuned model wins.
Alignment and PBO both learn a latent score from answers to "which is better?" with a pairwise random utility likelihood. Alignment then uses the score to reshape a distribution over text near a reference model; PBO uses it to find one best design. Most of what follows comes back to that difference.
Sources cited in Section 35.1 6
- Ouyang et al. (2022) Training language models to follow instructions with human feedback
- Christiano et al. (2017) Deep reinforcement learning from human preferences
- Coste et al. (2024) Reward Model Ensembles Help Mitigate Overoptimization
- Rafailov et al. (2023) Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Tang et al. (2024) Generalized Preference Optimization: A Unified Approach to Offline Alignment
- Azar et al. (2024) A General Theoretical Paradigm to Understand Learning from Human Preferences
35.2 Language models inside the preference loop #
Between 2023 and 2026 language models joined preference loops in five jobs: a prior or warm start, a simulated user, a translator from text to preference labels, a feature extractor, and a generator of candidates. The systems that worked best kept a classical probabilistic model in charge of uncertainty and of choosing queries, and let the language model work only at the interface; systems in which the language model itself acted as the optimizer fell behind classical dueling-bandit algorithms in strong regret.
35.2.1 Translators: from words to preference signals #
PEBOL (Austin et al., 2024a) elicits preferences in conversational recommendation with an independent Beta-distributed utility per item, the conjugate model of Section 5.2: a natural language inference model turns the user's "yes" or "no" about an aspect into a likelihood, and Thompson sampling or an upper confidence bound (Section 12.5, Section 12.4) decides what GPT-3.5 should ask next. After 10 turns, its mean reciprocal rank at 10 (the average of one over the rank of the user's target item in the top-10 list) was 0.27 on Yelp against 0.12 for a monolithic GPT-3.5 elicitor, 0.18 against 0.09 on MovieLens, and 0.17 against 0.11 on Recipe-MPR; the best monolithic baseline, Gemini-Pro, reached 0.17. The authors attribute the monolithic models' failure to over-exploitation, in severe cases asking the same question again. The users were 100 GPT-3.5 simulations per data set, told in advance which item they liked.
OPEN (Handa et al., 2024), a 2024 preprint, has a language model extract and rank features of a domain to initialize the prior of a linear Bradley-Terry utility, chooses pairwise queries by the expected information gain of Section 6.4 over a particle-filter posterior (a cloud of weighted samples), and has the language model rewrite each abstract comparison as a natural question. With people on the Prolific platform, recommending New York Times articles, it beat elicitation by the language model alone and by experimental design alone; without the rewriting, users found it markedly harder to express their preferences accurately, and mental demand was similar across methods. MAPLE (Mahmud et al., 2025) (AAAI 2025) turns natural-language feedback into samples or ranges for the linear weights of abstract concepts, with a Bradley-Terry likelihood over pairwise rankings of trajectories and Markov chain Monte Carlo (Section 17.5), evaluated with modeled humans.
LILO (Kobalczyk et al., 2026) (ICML 2026) is the closest to PBO in the strict sense. A language model translates free-text feedback and textual priors into pairwise labels for a probit pairwise Gaussian process over a composite utility , where gives an experiment's measured outcomes and the decision maker's utility over them, the setting of preference exploration (Lin et al., 2022), and qEUBO, the expected utility of the best option for queries of options (Section 19.4), selects which outcomes to have compared. Across 10 environments it beat pure language-model optimizers and PBO baselines, and in some settings one message from the decision maker matched or exceeded 8 to 16 pairwise comparisons; the problems had 5 to 8 dimensions, and the decision maker was simulated by Llama-3.3-70B. Its precursors include preregistered experiments with people in which questions generated by a language model elicited answers often more informative than prompts or labels users wrote themselves, with less reported effort (Li et al., 2025b); a NeurIPS 2023 workshop paper that chose questions by expected entropy reduction (Piriyakulkij et al., 2023); and, by LILO's first author, Bayesian experimental design over candidate solutions sampled from a language model (Kobalczyk et al., 2025) (ICLR 2025).
35.2.2 Priors, features, candidates, and the optimizer itself #
Priors and warm starts. Besides OPEN and MAPLE, Eichelbeck et al. (2026) (ICLR 2026) run PBO in an autoencoder's latent space with a prior a language model generates from interviews, with simulated users only. On the scalar side (Section 30.6), LLAMBO (Liu et al., 2024a) (ICLR 2024) uses a language model for warm starts, as a surrogate, and to sample candidates; a 2025 reproduction preprint found that warm starting "substantially improves early regret behaviour and reduces variance across runs" but that the surrogate is "weaker than GP or SMAC as a pure single task regressor" (SMAC uses a random forest), and its abstract attributes LLAMBO to "Daxberger et al. (2024)", which does not match the original authors (Rychert et al., 2025). LGBO (Yuan et al., 2026) (ICLR 2026) adds a language model's preferences about regions to the surrogate mean, with a worst-case guarantee, faster convergence when the preferences agree with the objective, and one wet-lab experiment, for a scalar objective. The evidence against is pointed: language-model agents in scalar BO performed no differently when their observed outcomes were replaced by randomly permuted labels, while linear bandits and Gaussian process optimization consistently won (Gupta et al., 2025) (EMNLP 2025 Findings), and the better starting point a language-model adviser suggested was a default configuration (Rodrigues et al., 2026), a 2026 preprint.
Feature extractors and surrogates. Language models help BO over molecules "only if they have been pretrained or finetuned with domain-specific" data (Kristiadi et al., 2024a) (ICML 2024); BO-ICL (Ramos et al., 2026), first posted in 2023 and published in 2026, found a near-optimal catalyst among 3,700 candidates within 6 iterations by in-context regression; and Ranković et al. (2026) (Nature Machine Intelligence 2026) train embeddings and a Gaussian process jointly, matching conventional BO with a median of 41% fewer iterations across 23 chemistry and materials tasks. All three use scalar feedback; with preference feedback, the instances of "language model embedding plus a probabilistic head" are APOHF and the method of Dwaracherla et al. (Section 35.3), neither with a Gaussian process.
Candidates and the optimizer itself. APPO (Li et al., 2026f) (CHI 2026) has a language model rewrite text-to-image prompts from a user's binary preferences, and PDO and Duel-Evolve (Section 35.3) let one language model both generate candidates and judge them. Xia et al. (2025) (ACL 2025 Findings) asked top language models to act as dueling-bandit algorithms in context. They quickly brought the best arm into the duels, giving low short-term weak regret (which counts the better of the two arms), but in strong regret (which counts both) "an optimality gap still exists" with classical algorithms; they "struggle to converge and consistently exploit even when explicitly prompted to do so". A hybrid, LEAD, wraps a classical algorithm around the language model and inherits its guarantees.
Table 35.1 sorts the main systems by whether they are PBO in the strict sense used in this book: a probabilistic surrogate that learns a latent utility over a continuous or structured design space from pairwise or ranked feedback, with an acquisition function choosing the queries.
| System | Role of the language model | Probabilistic model and query rule | Feedback | Strict PBO? | Who answered in the evaluation |
|---|---|---|---|---|---|
| PEBOL (RecSys 2024) | asks questions; an inference model turns answers into a likelihood | independent Beta-Bernoulli item utilities; Thompson sampling or UCB | yes or no | no: independent-arm bandit | 100 GPT-3.5 simulated users per data set |
| OPEN (2024 preprint) | extracts features, initializes the prior, rewrites questions | linear Bradley-Terry utility; particle filter; expected information gain | pairwise | partly: no GP | people on Prolific |
| MAPLE (AAAI 2025) | supplies weight priors, interprets language feedback | linear weights over concepts; Bradley-Terry; MCMC | pairwise rankings plus language | partly: as for OPEN | modeled humans |
| LILO (ICML 2026) | translates free text into pairwise labels | probit pairwise GP on a composite utility; qEUBO | language turned into pairwise labels | yes | Llama-3.3-70B simulated decision maker, 5 to 8 dimensions |
| Eichelbeck et al. (ICLR 2026) | generates the prior from interviews | PBO in an autoencoder latent space | pairwise | yes | simulated users only |
| LGBO (ICLR 2026) | states regional preferences | shift of the GP mean | scalar plus language-model preferences | no: scalar BO | benchmarks and one wet-lab experiment |
| APOHF (workshop paper) | frozen embeddings as features | neural network, Bradley-Terry; greedy plus UCB | pairwise | no: neural dueling bandit | simulated (validation accuracy, image similarity) |
| Duel-Evolve (workshop paper) | generates candidates and judges them | Bayesian Bradley-Terry; double Thompson sampling | the model's own pairwise judgments | no: test-time search | ground-truth benchmark accuracy |
Only LILO and the method of Eichelbeck et al. meet the strict definition, so calling all of these "PBO with language models" would overstate how much of this intersection Gaussian process preference models occupy.
35.2.3 Language for goals, comparisons for judgments #
If a language model can carry a person's words into the loop, should the person still be asked to compare? In the design studies of Chapter 32, natural language and explicit constraints gave the same optimization performance, with lower workload for language and a stronger sense of agency for constraints, and 90.9% of 187 natural-language requests described a desired outcome rather than a parameter value (Niwa et al., 2025). In a preprint, Peng et al. (2026) found that 20 people judging the same 600 pairs of generated interfaces agreed with one another at a Krippendorff's of only 0.25 (two of them made the same choice on 62.4% of pairs on average), yet for 12 new users, personalization from just 8 pairwise judgments beat every baseline, including the users' own written preferences, with an aggregate win rate of 60.35% (Section 32.4).
The model side points the same way. Language models inferring a user's preferences over several turns fall "far short" of normative Bayesian updating, and training them to imitate a Bayesian model improves this and generalizes (Qiu et al., 2026) (Nature Communications, January 2026); and on a simulated-user benchmark for helping users construct preferences, no frontier model exceeded 56% accuracy within 5 turns (Saracay et al., 2026) (COLM 2026). Together these results support a division of labor: language for goals, constraints, and priors; comparisons for fine judgments; and posterior updating left to an explicit probabilistic model (inference), as in the case for comparisons over ratings in Section 16.1.
35.2.4 Language models as simulated users and judges #
Many systems in this chapter were evaluated with a language model standing in for the person. The evidence from 2023 to 2026 points one way: close to people in aggregate, unreliable for individuals, and biased in systematic directions. In Table 35.2, Spearman correlation and Kendall's both measure how similarly two lists are ordered (1 for identical order, 0 for no relation), and agreement rates count how often two judges pick the same option.
| Study | Setting | In aggregate | For individuals, or biases |
|---|---|---|---|
| Zheng et al., NeurIPS 2023 Datasets and Benchmarks | GPT-4 judging MT-Bench and Chatbot Arena | agreement with people "over 80%", the same as between people | position bias, verbosity bias, self-enhancement bias |
| Dubois et al., NeurIPS 2023, AlpacaFarm | GPT-4 annotator with a single prompt | 65% agreement with people against 66% between people; method rankings correlate with those from human data at Spearman 0.98, at about 50 times lower cost | does not reproduce the variability of human labels or reward over-optimization; reproducing it needed a pool of simulated annotators and labels flipped at random with probability 0.25 |
| Wang et al., ACL 2024 | ChatGPT as evaluator | not reported | after swapping the order of responses, Vicuna-13B beat ChatGPT on 66 of 80 queries |
| Panickssery, Bowman, and Feng, NeurIPS 2024 | fine-tuning changes self-recognition | not reported | self-recognition ability correlates linearly with the strength of self-preference |
| Muldrew et al., ICML 2024 | 50 prompts labeled twice | GPT-4 agrees with itself more than 90% of the time | GPT-3.5-turbo about 60% |
| Kirk et al., 2026 preprint, PRISM-X | 530 people, each ranking four models; GPT-4o as simulated user, either ranking the person's transcripts or also holding the conversations | model scores fitted to simulated and to human rankings correlate at r = 0.99 and 0.98 | mean Kendall's τ with the person's own ranking of 0.22 (ranking only) and 0.11 (conversing), against 0.57 for people's own consistency; strong position effects (Figure 35.1); more sycophantic than people |
| Kuric, Demcak, and Krajcovic, 2026 preprint | 29 real design preference tests (n = 2073), 78 tasks | top-choice agreement 53%; GPT 5.2 reaches 65% with no significant gain in fidelity | simulated and real distributions differ significantly in 44% of tasks; single-persona simulation deviates significantly in 91%; temperature and top-p have no significant effect |
| Xu et al., ICML 2025 | judges on AlpacaEval | round-robin plus Bradley-Terry raises Spearman correlation with Chatbot Arena from 95.0% to 96.4%, Kendall from 82.1% to 86.3% | judgments are intransitive; rankings depend on the chosen baseline model |
| Shi et al., IJCNLP-AACL 2025 | 15 judges, over 150,000 instances | not reported | position bias correlates weakly with prompt length and is "strongly affected by the quality gap between solutions" |
| Yang et al., 2026 preprint | 20 language models, pairs of responses of equal quality | not reported | capability and low self-preference bias are "often uncorrelated, or even negatively correlated"; structured multi-dimensional evaluation cuts the bias by 31.5% on average |
| Chawla, Thompson, and Young, 2026 preprint | 7 models, 4 tasks | not reported | intransitivity cannot be explained by a single ordering under any monotone link; a mixture of several latent orderings fits better |
| Seshadri et al., 2026 preprint | simulated users in agent evaluation | not reported | agent success rates differ by up to 9 percentage points between simulators; fidelity is worse for some dialect groups |
| Zhu, Huang, and Sang, WWW 2024 workshop | simulated users in conversational recommendation | not reported | data leakage in the dialogue history and the simulator's replies inflates results; PEBOL relies on exactly this kind of evaluation |
The sources are, in order, Zheng et al. (2023), Dubois et al. (2023), Wang et al. (2024a), Panickssery et al. (2024), Muldrew et al. (2024), Kirk et al. (2026), Kuric et al. (2026), Xu et al. (2025a), Shi et al. (2025), Yang et al. (2026a), Chawla et al. (2026), Seshadri et al. (2026), and Zhu et al. (2024). Position bias is a preference for whichever response is shown first (or second); verbosity bias a preference for longer responses; self-enhancement or self-preference bias a judge's preference for text it generated itself; sycophancy a tendency to tell the user what they appear to want to hear. Figure 35.1 draws the clearest case, PRISM-X.
Things to try:
- With Simulated user set to Ranks transcripts only, compare the top bar of Agreement with people with the bar below it: near-perfect agreement about which model is better overall, and a per-person agreement of 0.22.
- Switch to Also holds the conversations. Per-person agreement halves to 0.11, and the share of trials in which the first-shown model wins rises from 33.8% to 44.9%, against 24.1% for people.
For PBO that uses a language model as its preference oracle, this evidence means the observation model is misspecified (inference). Language-model labels are faithful in aggregate and wrong for individuals (PRISM-X, Kuric et al.), too consistent (AlpacaFarm needed 25% random flips to reproduce human-like over-optimization), dependent on presentation order (Wang et al.), and biased toward the judge's own text (Panickssery et al.). A probit or Bradley-Terry likelihood treats every deviation as independent noise with a fixed variance, but most of these deviations are bias, and a position bias that varies with the quality gap (Shi et al.) makes the noise variance depend on the utility difference, which a fixed noise scale cannot express (Section 27.2 lists heteroscedastic models). The most direct remedy is the balanced position calibration of Wang et al.: average the judgments from both orders. NAOD (Du et al., 2026), a preprint posted on September 30, 2026, is the first method we found that puts judge bias into the design of queries: arguing that active acquisition exposes judge bias that remains after calibration, it lowered average proxy policy regret by 29.1% on Chatbot Arena data with 17 judges, against a matched design that targets information alone, and showed that representation error "can reverse an oracle design advantage". LILO does not model judge bias.
As of September 2026 we found no study that runs the same PBO loop with human comparisons and with language-model comparisons and reports regret or sample efficiency for both, and none that measures how position bias propagates into acquisition decisions (for example, whether averaging over both orders changes the sequence of selected queries). Until such a study exists, a PBO result obtained with a language-model oracle is a result about that oracle; the minimal check is to rerun a few sessions with people and with both presentation orders, and to report how the selected queries change (inference). The gap is not peculiar to language models: in a retinal-implant study within PBO, only about 50% of human choices agreed with a simulated agent that was not a language model, yet 16 of 17 participants preferred the optimized result (Schoinas et al., 2025) (Section 33.5, Section 31.7).
35.2.5 The pattern #
From PEBOL in 2023 to LILO in 2026, the model of uncertainty stayed classical and small (PEBOL's Beta-Bernoulli model, the linear Bradley-Terry models of OPEN and MAPLE, LILO's pairwise Gaussian process), and the language model worked only at the interfaces: features, prior initialization, wording of questions, and translation of labels. Every approach that handed query selection to the language model fell behind classical algorithms (inference).
Sources cited in Section 35.2 38
- Austin et al. (2024a) Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation
- Handa et al. (2024) Bayesian Preference Elicitation with Language Models
- Mahmud et al. (2025) MAPLE: A Framework for Active Preference Learning Guided by Large Language Models
- Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Li et al. (2025b) Eliciting Human Preferences with Language Models
- Piriyakulkij et al. (2023) Active Preference Inference using Language Models and Probabilistic Reasoning
- Kobalczyk et al. (2025) Active Task Disambiguation with LLMs
- Eichelbeck et al. (2026) Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent Space
- Liu et al. (2024a) Large Language Models to Enhance Bayesian Optimization
- Rychert et al. (2025) Reproducibility Study of Large Language Model Bayesian Optimization
- Yuan et al. (2026) Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery
- Gupta et al. (2025) LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
- Rodrigues et al. (2026) When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model
- Kristiadi et al. (2024a) A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
- Ramos et al. (2026) Bayesian Optimization of Catalysis With In-Context Learning
- Ranković et al. (2026) Large language models as uncertainty-calibrated optimizers for experimental discovery
- Li et al. (2026f) Preference-Guided Prompt Optimization for Text-to-Image Generation
- Xia et al. (2025) Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Peng et al. (2026) Efficient Personalization of Generative User Interfaces
- Qiu et al. (2026) Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
- Saracay et al. (2026) Beyond expert users: agents should help users construct preferences, not just elicit them
- Zheng et al. (2023) Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Dubois et al. (2023) AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
- Wang et al. (2024a) Large Language Models are not Fair Evaluators
- Panickssery et al. (2024) LLM Evaluators Recognize and Favor Their Own Generations
- Muldrew et al. (2024) Active Preference Learning for Large Language Models
- Kirk et al. (2026) PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
- Kuric et al. (2026) Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design?
- Xu et al. (2025a) Investigating Non-Transitivity in LLM-as-a-Judge
- Shi et al. (2025) Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
- Yang et al. (2026a) Quantifying and Mitigating Self-Preference Bias of LLM Judges
- Chawla et al. (2026) Multiple latent orderings better predict language model preferences
- Seshadri et al. (2026) Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
- Zhu et al. (2024) How Reliable is Your Simulator? Analysis on the Limitations of Current LLM-based User Simulators for Conversational Recommendation
- Du et al. (2026) Optimal Design for Active Preference Learning with Biased LLM Judges
- Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
35.3 Preference methods for language-model problems #
The second direction applies tools built for dueling bandits and PBO to problems of language models: optimizing prompts, choosing among outputs and models, choosing which comparisons to collect for training, and exploring online during alignment. The evidence is mixed in a way that turns out to be informative, and it resolves once one asks what each query is for.
35.3.1 Prompt optimization #
APOHF (Lin et al., 2024b) trains a neural network on frozen prompt embeddings with a Bradley-Terry likelihood and pairs the greedy maximizer with the maximizer of an upper confidence bound, the recipe of linear dueling bandits, against baselines that included an ensemble with double Thompson sampling, a dueling-bandit rule that draws two independent posterior samples and pairs their maximizers (Section 21.2). It handles only pairs, and all its "human" feedback was simulated: validation accuracy on 30 instruction tasks and similarity to a target image for text-to-image tasks. It was an ICML 2024 workshop oral, was not accepted at ICLR 2025, and has an ACL Rolling Review submission record from May 2026 with no formal acceptance. Kayal et al. (2025) (ICML 2025) use it as the motivating application of "Bayesian optimization from human feedback". PDO (Wu et al., 2026) (ACL 2026 Findings) schedules duels between prompts with double Thompson sampling under a fixed budget of language-model judgments and uses the winners to guide mutation, finding prompts better than label-free baselines on BIG-bench Hard and MS MARCO. Among the prompt optimizers for text we found, none was evaluated with people; the study of APPO with people concerns images. (Ji, He, and Gu also call an unrelated active-query algorithm APPO.)
35.3.2 Choosing among outputs and among models #
Zhang et al. (2024a) (ICML 2024) replace pointwise scoring of reasoning steps with pairwise comparisons by a language model, handling the noise with ensembles and dueling-bandit variants. Chatbot Arena (Chiang et al., 2024), the public leaderboard built from crowd votes between anonymous models, chooses the next pair in proportion to the expected reduction in the width of the Bradley-Terry confidence intervals; in simulations fitted to 213,576 held-out votes, random sampling needed 54% more data to reach a precision of 0.2 on the win-rate matrix, but only 5% more to reach 0.3 on the scores. Duel-Evolve (Karlekar et al., 2026), an ICLR 2026 workshop poster, fits a Bayesian Bradley-Terry model to a language model's judgments of its own candidates, allocates comparisons by double Thompson sampling, and picks parents for the next generation in evolutionary fashion, with no reward model or labels; it scored 20 percentage points above comparable iterative methods on MathBench and 12 or more on LiveCodeBench, against ground truth. Model routing has been cast as a contextual dueling bandit solved with Feel-Good Thompson sampling, a variant with an optimism term, with lower cumulative regret on RouterBench and MixInstruct (Chiang et al., 2025) (a preprint not accepted at ICLR 2026); Gharat et al. (2026) (NeurIPS 2026) study best-arm identification under dueling feedback with heterogeneous query costs, assuming a Condorcet winner (an option that beats every other with probability above one half) and checking that assumption on real data; CUPID (Nguyen et al., 2026) (ICML 2026) chooses which two models a user should compare, with a study with people; and T-POP (Qu et al., 2026) (ICML 2026) learns one user's reward online with a dueling bandit to steer the decoding of a frozen model.
35.3.3 Active preference data for reward models and DPO #
Choosing which comparisons to collect for training has the most evidence. Muldrew et al. (2024) (ICML 2024) favor pairs on which DPO's implicit preference model is confident but wrong, and with GPT-4 as the oracle improved win rate by 1% to 6% on average. BAL-PM (Melo et al., 2024) (NeurIPS 2024) observes that naive estimates of epistemic uncertainty (uncertainty from lack of data, as opposed to noise in the answers) select redundant samples, adds the entropy of the already-acquired prompts in the model's feature space, and needed 33% to 68% fewer labels on two human preference data sets. Mehta et al. (2025) (COLM 2025) reduce dueling feedback to a contextual Borda function, the probability that an action beats a uniformly random one, prove a polynomial regret bound, and report that their methods can improve performance by over 13% relative to baselines under a limited budget, with uncertainty from Monte Carlo dropout (keeping dropout on at prediction time). Das et al. (2025) (ECML-PKDD 2025) prove that sampling contexts uniformly can leave a constant suboptimality gap under a small budget, give a lower bound of for dimensions and rounds, and show that APO, which picks the most uncertain context, matches it up to logarithmic and nonlinear factors. Ji et al. (2024) (TMLR) prove a regret of and a query complexity of , where is the gap between the best and second-best actions; their ADPO matches DPO with about half as many queries.
A branch based on optimal experimental design uses the Fisher information, the expected curvature of the log-likelihood, whose inverse approximates the covariance of the fitted parameters; a D-optimal design maximizes its determinant, shrinking the uncertainty ellipsoid. It includes an offline design with a lower bound matching up to constant and logarithmic factors (Scheid et al., 2024), D-optimal feedback for a DPO linearized at the last layer (Kveton et al., 2025) (both preprints), a Fisher criterion at the reward model's last layer, with the finding that comparisons across prompts raise labeling efficiency (Shen et al., 2025a) (ICML 2025), a log-determinant choice of negative examples for the Plackett-Luce model (Surana et al., 2026) (a 2026 preprint), and PBO-style acquisition in RLHF with a last-layer Laplace approximation (Section 17.2) (Cercola et al., 2026a) (published in 2026).
ActiveUltraFeedback (Melikidze et al., 2026) (ICML 2026) is the most systematic comparison so far: a contextual dueling bandit over responses from 12 families of open models, an ensemble on a frozen backbone, and a language-model judge's Likert ratings standing in for people. It compared random selection, two heuristics, three classical dueling-bandit rules (infomax, the pair whose choice probability the ensemble disagrees on most; double Thompson sampling; and MaxMinLCB), and two new rules that seek a large quality gap: double reverse Thompson sampling, which pairs the maximizer and minimizer of one posterior sample, and DeltaUCB. Models fine-tuned on 5,000 to 10,000 samples chosen by the new rules outperformed models trained on 60,000 samples chosen by random selection, by UltraFeedback, or by the dueling-bandit rules. For training a reward model, though, random selection was strong, and the authors conclude that there diversity is preferable to a large quality gap.
35.3.4 Thompson sampling for online alignment #
Dwaracherla et al. (2024) (ICML 2024) brought double Thompson sampling and epistemic neural networks, which output a family of predictions indexed by a random input so that its spread expresses what the network does not know, into feedback collection for language models. Queries were a prompt and two Gemini Nano responses; the "human" was a simulator choosing by the Bradley-Terry model from a reward model built on Gemini Pro and fitted to Anthropic's helpfulness and harmlessness data; and the policy was approximated by best-of- (keep the highest-scoring of responses). Double Thompson sampling beat passive querying, Boltzmann exploration, and infomax, which did well early but then fell far behind, which the authors attribute to its seeking information whether or not it is useful; it reached passive querying's performance with an order of magnitude less data. Asghari et al. (2026), a 2026 preprint, scale this to a 9B Gemma policy, with an epistemic reward network adding fewer than 5% to the parameters, pairs chosen by the variance of the choice probability across ensemble particles (information-directed selection), and a small positive bias on each reinforcement signal, which the authors call an affirmative nudge. With fewer than 20,000 labels it matched what offline RLHF reached with 200,000. The factor of 1,000 in the paper is an extrapolation of curves, the ablations do not separate the selection rule from the nudge, and there is no independent replication.
The theory is well developed: a near-optimal trade-off between regret and queries for trajectory preferences (Wu and Sun, 2024) (ICLR 2024; Section 36.1); FGTS.CDB (Li et al., 2024b) (ICML 2024), the first posterior sampling algorithm for linear contextual dueling bandits, with near minimax optimal regret ; neural-tangent-kernel versions of the upper confidence bound and Thompson sampling (Verma et al., 2025) (ICLR 2025); and an bound for Thompson sampling in online RLHF with general function approximation (Feng and Fu, 2025) (a 2025 preprint, without experiments).
Two applied versions followed. SEA (Liu et al., 2024b) frames alignment as a contextual dueling bandit solved with Thompson sampling on models of 1B to 6.9B parameters with DPO, IPO, and SLiC, its preferences coming from a scalar reward model (Skywork-Reward-Llama-3.1-8B) plus one experiment with GPT-4o-mini as judge; it is a NeurIPS 2024 workshop poster, not accepted at ICLR 2025, NeurIPS 2025, or ICLR 2026. warmPref-PS (Agnihotri et al., 2024), a preprint not accepted at TMLR, warm-starts posterior sampling with offline preferences from an expert of unknown competence.
A parallel branch replaces posterior sampling with exploration bonuses: XPO (Xie et al., 2025) (ICLR 2025), value-incentivized preference optimization (Cen et al., 2025) (ICLR 2025), self-exploring language models (Zhang et al., 2024b) (TMLR), and a count-based bonus added to the DPO objective (Bai et al., 2025) (ICLR 2025). Under KL or -divergence regularization (a family that contains the KL) these bonuses "unintentionally" bias exploration toward the reference model's high-probability regions (Li et al., 2026d) (ICLR 2026); the sample complexity of all existing online RLHF algorithms grows exponentially with the scale of the reward (Chen et al., 2025a) (NeurIPS 2025); and iterative Nash preference optimization (Section 35.4.4) without explicit exploration can depend exponentially on the KL parameter (Nan et al., 2026) (a 2026 preprint).
35.3.5 Random is hard to beat #
Oh et al. (2026b), an ICLR 2026 workshop paper (I Can't Believe It's Not Better), compared uncertainty-based active preference learning with random selection in online DPO, across harmlessness, helpfulness, and instruction following, with a reward model and a language-model judge as proxies. Active selection "yields negligible improvements in proxy win-rates compared to Random"; while proxy win rate rose, capability on standard benchmarks fell, and active selection neither prevented that nor clearly reduced variance. The authors point to the strong prior from web-scale pretraining and the "cheap diversity" of random on-policy samples. Other results qualify the positive ones: naive uncertainty selects redundant samples (BAL-PM), infomax fell behind after the early phase (Dwaracherla et al.), quality-gap rules beat the classical dueling rules (ActiveUltraFeedback), and Chatbot Arena's savings were 54% for one target and 5% for another. And the positive results simulate people with a reward model (Dwaracherla et al., Asghari et al.), use accuracy or similarity as the utility (APOHF), or use language-model judges (Muldrew et al., ActiveUltraFeedback, PDO, Duel-Evolve); BAL-PM's savings come from human preference data sets, and Chatbot Arena's from simulations fitted to real votes.
35.3.6 What the query is for #
The results can be reconciled once the purpose of a query is separated (inference). The acquisition rules of dueling bandits and PBO (double Thompson sampling, MaxMinLCB, qEUBO, infomax) were designed to find or confirm the best option. Alignment data need pairs that teach a policy or a reward model: DPO training benefits from a large quality gap (ActiveUltraFeedback), reward-model training from diversity (the random-selection results of BAL-PM and ActiveUltraFeedback), and offline methods from coverage of the policy's own distribution (Oh et al., and Song et al. in Section 35.4.3). Using a find-the-best rule unchanged answers a different question from the one being asked. Two further differences matter: alignment picks from a small, policy-dependent candidate set rather than a whole design space, so much of the disagreement between "active" and "random" may come from how diverse that set already is (inference); and Asghari et al. have an explicit posterior over rewards, whereas Oh et al. have only the implicit reward of a directly optimized policy.
As of September 2026 we found no study that compares double Thompson sampling, information-directed selection, quality-gap selection, and random selection under the same budget, with human labels, for both reward-model RLHF and DPO. Until one does, a team collecting preference data should keep a random arm as its baseline and choose the active rule by what the data are for (inference).
Sources cited in Section 35.3 37
- Lin et al. (2024b) Prompt Optimization with Human Feedback
- Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
- Wu et al. (2026) LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
- Zhang et al. (2024a) Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
- Chiang et al. (2024) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Karlekar et al. (2026) Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
- Chiang et al. (2025) LLM Routing with Dueling Feedback
- Gharat et al. (2026) Cost-Aware Best-LLM Identification using Dueling Feedback
- Nguyen et al. (2026) CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM
- Qu et al. (2026) T-POP: Test-Time Personalization with Online Preference Feedback
- Muldrew et al. (2024) Active Preference Learning for Large Language Models
- Melo et al. (2024) Deep Bayesian Active Learning for Preference Modeling in Large Language Models
- Mehta et al. (2025) Sample Efficient Preference Alignment in LLMs via Active Exploration
- Das et al. (2025) Active Preference Optimization for Sample Efficient RLHF
- Ji et al. (2024) Reinforcement Learning from Human Feedback with Active Queries
- Scheid et al. (2024) Optimal Design for Reward Modeling in RLHF
- Kveton et al. (2025) Active Learning for Direct Preference Optimization
- Shen et al. (2025a) Active Reward Modeling: Adaptive Preference Labeling for Large Language Model Alignment
- Surana et al. (2026) MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization
- Cercola et al. (2026a) Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference
- Melikidze et al. (2026) ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
- Dwaracherla et al. (2024) Efficient Exploration for LLMs
- Asghari et al. (2026) Efficient Exploration at Scale
- Wu and Sun (2024) Making RL with Preference-based Feedback Efficient via Randomization
- Li et al. (2024b) Feel-Good Thompson Sampling for Contextual Dueling Bandits
- Verma et al. (2025) Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
- Feng and Fu (2025) Thompson Sampling in Online RLHF with General Function Approximation
- Liu et al. (2024b) Sample-Efficient Alignment for LLMs
- Agnihotri et al. (2024) Online Bandit Learning with Offline Preference Data for Improved RLHF
- Xie et al. (2025) Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
- Cen et al. (2025) Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
- Zhang et al. (2024b) Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
- Bai et al. (2025) Online Preference Alignment for Language Models via Count-based Exploration
- Li et al. (2026d) General Exploratory Bonus for Optimistic Exploration in RLHF
- Chen et al. (2025a) Avoiding $\mathbfexp(R_max)$ scaling in RLHF through Preference-based Exploration
- Nan et al. (2026) Efficient Exploration for Iterative Nash Preference Optimization
- Oh et al. (2026b) Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
35.4 A shared likelihood and seven broken analogies #
The two fields share an observation model and, for the online problem, a decision-theoretic framing, so an acquisition rule from one field can run in the other almost unchanged. That shared core is easy to over-read; this section separates it from seven places where the analogy breaks.
35.4.1 What is shared #
The likelihood. Zhu et al. (2023) (ICML 2023) prove that with linear
rewards the maximum likelihood estimate converges under both the
Bradley-Terry-Luce and Plackett-Luce models but that a policy trained on it can
fail, whereas a pessimistic estimate works under a coverage assumption; that
for rankings of items the full Plackett-Luce estimate is "asymptotically
more efficient" than splitting the ranking into pairs; and that RLHF and
maximum-entropy inverse reinforcement learning share one analysis.
Sun et al. (2025) (ICLR 2025) give convergence rates for Bradley-Terry reward
models on deep embeddings and argue that Bradley-Terry is "not a necessary
choice", since downstream optimization needs only an order-consistent reward.
Tang et al. (2024) (ICML 2024) show that DPO, IPO, and SLiC differ only in
their convex loss, which corresponds to the choice of link in PBO. On the PBO
side, BoTorch's PairwiseGP (Meta Platforms, Inc., 2026h) (software documentation)
and LILO use the probit link of Thurstone (Section 16.3, Section 27.1),
while Kayal et al. and PF-TS (Lazzaro et al., 2026) (AISTATS 2026) analyze the
Bradley-Terry-Luce model.
The online decision problem. Xiong et al. (2024) (ICML 2024) formalize RLHF as a "reverse-KL regularized contextual bandit" with finite-sample guarantees, and Ji et al., Mehta et al., and SEA adopt the contextual dueling bandit form. Regret guarantees exist on both sides (Kayal et al. and PF-TS for PBO, Section 29.3; Wu and Sun, Ji et al., and APO for RLHF), and acquisition rules transfer almost unchanged.
35.4.2 Where the analogy breaks #
The objective. By Equation (35.5), the optimum of the KL-regularized objective lets the policy network represent both the language model and the implicit reward (Rafailov et al., 2023). Korbak et al. (2022) (EMNLP 2022 Findings) observe that plain reinforcement learning fine-tuning leads to "distribution collapse", whereas KL-regularized reinforcement learning is "equivalent to variational inference", approximating a Bayesian posterior that updates the prior language model with the reward as evidence. PBO returns the maximizer of a latent utility; alignment returns the reference policy tilted exponentially by the reward, and measures regret against that tilted policy. Won et al. (2025), a 2025 preprint, argue that when preferences encode the information needed to update a reference policy into a target, the log-ratio reward is the only reasonable choice.
The implicit reward as a utility estimate. Reading DPO's implicit reward like a Gaussian process posterior mean is not reliable: it fits the training data as well as an explicit reward model but is 3% less accurate on average, and up to 7% less, across five out-of-distribution settings (Lin et al., 2024a) (EMNLP 2024 Findings), and most preference-tuned models rank the pairs of common preference data sets correctly "less than 60%" of the time, with the DPO objective ill-suited to correcting even mild ranking errors of the reference model (Chen et al., 2024) (NeurIPS 2024).
Strong preferences. When preferences are close to deterministic, Bradley-Terry reward differences diverge and the KL regularization grows ever weaker, so DPO overfits, especially when each pair appears only a few times (Azar et al., 2024) (AISTATS 2024). In PBO the Gaussian process prior constrains the latent utility directly, so the problem does not arise in the same form (inference). Table 35.3 lays out all nine aspects.
| Aspect | PBO assumes | Preference-based alignment assumes | Holds? | Why, and the evidence |
|---|---|---|---|---|
| Observation model | probit link (PairwiseGP, LILO) or logistic Bradley-Terry (Kayal et al., PF-TS) |
logistic Bradley-Terry; Plackett-Luce for lists; DPO, IPO, and SLiC correspond to logistic, squared, and hinge losses | holds | both are pairwise random utility likelihoods; the links differ only in the noise distribution (Zhu et al. 2023; Tang et al. 2024; Rafailov et al. 2023) |
| Online decision problem | a dueling bandit with cumulative or simple regret | a contextual dueling bandit with reverse-KL regularization | holds | the same acquisition rules apply (Xiong et al. 2024; Ji et al.; Mehta et al. 2025) |
| 1. Objective | the maximizer of the latent utility | the reference policy tilted exponentially by the reward | breaks | regret is measured against a tilted reference policy (Rafailov et al. 2023; Korbak et al. 2022) |
| 2. Sample complexity | conditional upper bounds: preference feedback of the same order as order-optimal scalar feedback (Kayal et al.) | rather than under KL regularization; existing online algorithms exponential in the reward scale; iterative Nash methods can be exponential in the KL parameter | breaks | the KL term to a reference policy, not the likelihood, drives the rates (Zhao et al. 2025; Chen et al. 2025; Nan et al. 2026) |
| 3. Strong preferences | the GP prior constrains the latent utility | near-deterministic preferences make reward differences diverge and the KL constraint loses force | breaks | nothing like a prior on the reward holds the fit in place (Azar et al. 2024) |
| 4. Utility estimate | a posterior mean and variance | DPO's implicit reward has no uncertainty; 3% less accurate out of distribution on average; ranking accuracy mostly below 60% | breaks | the implicit reward is a by-product of fitting a policy (Lin et al. 2024; Chen et al. 2024) |
| 5. Query distribution | any pair anywhere in the design space | candidates sampled from the policy or a model pool; offline methods need global coverage; exploration bonuses drift toward the reference model's high-probability region | breaks | what can be compared is limited to what the policy generates (Song et al. 2024; Li et al. 2026) |
| 6. Scale | tens to hundreds of comparisons, about 2 to 20 dimensions, Laplace or expectation propagation | to comparisons over sequences; ensembles, dropout, or Laplace | breaks | a summary across the studies in this chapter (inference) |
| 7. Respondents | one decision maker with a consistent utility | many annotators or language-model judges | breaks | a single reward fitted to many people raises problems of social choice (Gölz et al. 2025; Siththaranjan et al. 2024; Chidambaram et al. 2026) |
Three of these differences, sample size, heterogeneity of the respondents, and where the queries come from, can be seen by fitting the same likelihood in both worlds, and Figure 35.2 also shows the first, what each world returns.
Things to try:
- Start in the one-person world and press New draw a few times. With 50 comparisons the band is wide and the returned design moves around the person's favorite near 0.75. Drag to 1,000,000: the band collapses onto the person's utility.
- Raise the second group's share to 0.4. The band stays thin, but the fitted reward matches neither group, and the error no longer shrinks as grows: one reward fitted to two groups is a compromise (Section 35.4.4).
- Switch Where pairs come from to From the policy. Outside the region the policy covers, the band stays wide however large grows: no amount of data reaches options the policy never generates.
- Switch the world to Many annotators. The orange curve, the reference policy reweighted by , can only reweight options the policy produces and so never reaches group A's favorite near 0.75. Back in One person with pairs from the policy and , PBO restricted to someone else's candidates stops at the edge of what the policy produces (Section 35.3.6).
35.4.3 Sample complexity and coverage #
Write for how close to optimal the learned policy must be. A sample complexity of means that halving needs four times the data; means twice. Zhao et al. (2025) (NeurIPS 2025) point out that earlier analyses of KL-regularized RLHF gave the same as the unregularized problem, and that a sharper analysis gives . Song et al. (2024) (NeurIPS 2024) prove that a global coverage condition, roughly that the data contain every response the optimal policy might produce, is necessary and sufficient for DPO-like offline contrastive methods to reach the optimal policy, while online reinforcement learning needs only partial coverage; step 3 above shows what a lack of coverage looks like. The main difference between the theories of PBO and RLHF therefore lies in the KL constraint to a reference policy and in the policy parameterization, not in the preference likelihood (inference).
35.4.4 Many annotators: from optimization to social choice #
The difference in respondents pushes alignment toward social choice, the study of how to combine many people's preferences into one decision (Section 40.7), which PBO has barely started on (Section 20.5). With cyclic preferences there may be no Condorcet winner. The Borda count scores each option by how often it beats an opponent drawn at random; a von Neumann winner is a probability distribution over options that beats or ties every single option in expectation, and exists even when preferences cycle; a Nash equilibrium of a two-player game is a pair of strategies neither player can improve on alone. Nash learning from human feedback (Munos et al., 2024) (ICML 2024) seeks a policy preferred to any opponent, the Nash equilibrium of a two-player constant-sum game, and with arbitrary preferences the target of preference-based reinforcement learning becomes the von Neumann winner (Wang et al., 2023a) (NeurIPS 2023). The von Neumann winner is also a solution concept of the dueling-bandit literature, so dueling-bandit results on intransitive preferences (Section 21.1) are the natural counterpart of Nash learning on the PBO side (inference); we did not check whether a paper from 2025 or 2026 makes this connection formally.
Gölz et al. (2025) (NeurIPS 2025) measure alignment methods by their distortion, the worst-case ratio between the best achievable average utility and that of the learned policy, and prove that Nash learning achieves minimax-optimal distortion while RLHF and DPO can have exponential or unbounded distortion in the full setting; Oko et al. (2026) (ICML 2026) then prove that under reward clipping the exponential degradation comes from a mismatch between the preference data and the reference policy, not from the algorithm. The dispute is open. What a single fitted reward does with many people has also been characterized: with unobserved context, such as which annotator answered, a single learned utility implicitly aggregates by Borda count (Siththaranjan et al., 2024) (ICLR 2024), as does the Bradley-Terry-Luce loss (An et al., 2026) (a 2026 preprint), which is the compromise curve of step 2 above; with limited data per user, binary comparisons cannot identify latent user types whereas rankings of three or more items can (Chidambaram et al., 2026) (AISTATS 2026); and under Bayesian marginalization or KL-robust optimization the effective reward of the KL-regularized objective has a closed form, optimistic in the Bayesian branch and pessimistic in the robust one (Hahami et al., 2026) (a 2026 preprint). The same reward posterior should then be used optimistically when choosing queries and pessimistically when choosing what to deploy, and the acquisition functions of PBO handle only the first (inference). Section 41.3 takes up what alignment to many people can mean.
Sources cited in Section 35.4 22
- Zhu et al. (2023) Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons
- Sun et al. (2025) Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
- Tang et al. (2024) Generalized Preference Optimization: A Unified Approach to Offline Alignment
- Meta Platforms, Inc. (2026h) BoTorch PairwiseGP source code pairwise_gp.py
- Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
- Xiong et al. (2024) Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
- Rafailov et al. (2023) Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Korbak et al. (2022) RL with KL penalties is better viewed as Bayesian inference
- Won et al. (2025) Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
- Lin et al. (2024a) On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
- Chen et al. (2024) Preference Learning Algorithms Do Not Learn Preference Rankings
- Azar et al. (2024) A General Theoretical Paradigm to Understand Learning from Human Preferences
- Zhao et al. (2025) Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
- Song et al. (2024) The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
- Munos et al. (2024) Nash Learning from Human Feedback
- Wang et al. (2023a) Is RLHF More Difficult than Standard RL?
- Gölz et al. (2025) Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
- Oko et al. (2026) Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner
- Siththaranjan et al. (2024) Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
- An et al. (2026) Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences
- Chidambaram et al. (2026) Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
- Hahami et al. (2026) A Unifying Lens on Reward Uncertainty in RLHF
35.5 Common claims, checked #
Table 35.4 checks common claims about how PBO relates to alignment.
| Claim | What the evidence says |
|---|---|
| PBO and RLHF share the Bradley-Terry model. | Partly. Both use pairwise random utility likelihoods, but the default PBO implementation, PairwiseGP, uses the probit link. Logistic forms appear in the original PBO paper of González et al. (González et al., 2017), in POP-BO (Xu et al., 2024b) (ICML 2024), in MaxMinLCB, and in MR-LPF (the algorithm of Kayal et al.). |
| DPO shares a probit or Bradley-Terry likelihood with PBO. | "Probit" is wrong. The DPO loss is the logistic Bradley-Terry log-likelihood, with a Plackett-Luce version in the appendix. |
| Bradley-Terry is the Rosetta stone connecting BO, RLHF, and DPO; DPO's implicit reward is conceptually like a Gaussian process posterior. | True at the level of the likelihood, and only an explanatory analogy for the implicit reward, which carries no uncertainty (Table 35.3). |
| RLHF scales but is query-inefficient; PBO is query-efficient but does not scale; a hybrid gets the best of both. | The direction holds, with qualifications. Hybrids exist (an ensemble posterior with double Thompson sampling or information-directed selection), but their posteriors are neural ensembles, not Gaussian processes, and the same kind of method gave negligible gains over random selection in online DPO. We found no work that uses a Gaussian process or kernel surrogate as the reward model in an RLHF loop at language-model scale. |
| Bringing PBO's uncertainty into RLHF, and warm-starting PBO with priors from language models, are open challenges. | Both have been pursued since 2024: the first by Dwaracherla et al., BAL-PM, and Cercola et al., with positive and negative results; the second by OPEN, MAPLE, and Eichelbeck et al., all but OPEN with simulated users. |
| Ji et al.'s RLHF with active queries appeared at NeurIPS. | The content is right (ADPO reaches DPO's level with about half the queries), but the venue is TMLR. |
| RLHF reward models are "notoriously poorly calibrated". | We could not substantiate this as stated, and it usually comes without a citation. Thies et al. (2026a) (ICML 2026) found that the calibration of RLHF reward models "correlates strongly but not perfectly with benchmark accuracy". |
| PBO uses about 10 to 1,000 comparisons and RLHF about 50,000 to 500,000 (or: tens to one or two hundred, against thousands to millions). | The contrast in order of magnitude holds, but neither range has a source. Studies of PBO with people use tens to one or two hundred comparisons; simulated high-dimensional PBO uses about 10 to 15 per dimension, about 1,500 at 102 dimensions (Menn et al., 2026a) (a 2026 preprint); alignment uses to . |
| LiPO shows that the gains from listwise preference data grow monotonically with list length. | Partly. Only the LiPO-λ loss benefits monotonically from longer lists; a Plackett-Luce form of DPO does not (Liu et al., 2025a) (NAACL 2025). |
| PEBOL raises MAP@10 by up to 131% after 10 turns. | That figure (mean average precision at 10) appears only in the first arXiv version (Austin et al., 2024b). The RecSys version reports mean reciprocal rank at 10 instead: at most 0.27, against 0.17 for the best monolithic baseline. |
Sources cited in Section 35.5 6
- González et al. (2017) Preferential Bayesian Optimization
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Thies et al. (2026a) Calibrated Preference Learning: The Case of Label Ranking
- Menn et al. (2026a) Local Preferential Bayesian Optimization
- Liu et al. (2025a) LiPO: Listwise Preference Optimization through Learning-to-Rank
- Austin et al. (2024b) Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation
35.6 What flows back #
The main direction of transfer has been from bandits and PBO into alignment: double Thompson sampling, epistemic neural networks, information-directed sampling, kernelized dueling bandits, and uncertainty-aware reward models. Transfer back has so far been mostly framing, applications, and evaluation. MaxMinLCB (Pásztor et al., 2024) (NeurIPS 2024), a kernelized preference bandit that casts the choice of a pair as a zero-sum Stackelberg game (one player commits first and the other responds), motivates itself by noting that such a model "has been employed in systems for fine-tuning large language models", and its group then published RewardUQ (Yang et al., 2026b), a EurIPS 2025 workshop paper, and ActiveUltraFeedback. Kayal et al. use prompt optimization as their application; PF-TS draws its two competitors symmetrically, which its authors argue is often desirable, and sometimes necessary, with human evaluators or language-model judges; and the multi-user dueling bandit of Ahmed and Ghasemi (2026) (TMLR) opens with the unfairness to minority groups of training on average preferences in language-model fine-tuning, and proves a lower bound on regret for fairness across users of with arms. Ax 1.2.3 added an abstraction for language-model messages, and 1.3.0 language-in-the-loop labeling trials and qEUBO scheduling for PBO (Meta Platforms, Inc., 2026l) (software documentation). Not everything that looks like a transfer is one: LILO extends BOPE (Lin et al., 2022) (AISTATS 2022), which predates the alignment wave, and the abstracts of PABBO (Zhang et al., 2025a) (ICLR 2025) and POP-BO do not frame their problems in terms of language models.
35.6.1 Three results ready to transfer #
Three results from alignment could be used in PBO directly. We found no PBO paper that takes DPO, LiPO, PAL, or an uncertainty-aware reward model as the source of a method component, though we did not systematically scan the reference lists of PBO papers from 2024 to 2026 (inference).
Rankings carry more than pairs. Zhu et al. proved that the full Plackett-Luce estimate beats splitting rankings into pairs, and Chidambaram et al. that rankings of three items identify latent user types that pairs cannot. Katkuri et al. (2026), a 2026 preprint, learn rewards with Plackett-Luce from a vision-language model's rankings and match or beat pairwise Bradley-Terry and RL-VLM-F (Wang et al., 2024b) (ICML 2024) on Meta-World; LiPO (Liu et al., 2025a) applies learning-to-rank losses (Section 36.6) to listwise preference optimization; and GraphDPO (Liu et al., 2026a), a 2026 preprint, generalizes preferences to a graph. The PBO interfaces that ask for choices among several options, GimmBO (Liu et al., 2026b) (SIGGRAPH North America 2026) and MultiBO (Rajagopalan et al., 2026) (ICML 2026), adopted lists for reasons of interface design (Section 20.1) and cite none of these formal arguments.
Separating noise from ignorance. Several alignment papers split reward uncertainty into an aleatoric part, the irreducible noise in the answers, and an epistemic part, uncertainty from lack of data; Lou et al. (2024), a 2024 preprint, use a probabilistic value head for the first and ensemble disagreement for the second. Others quantify reward uncertainty to mitigate over-optimization, with reward-model ensembles (Coste et al., 2024) (ICLR 2024) or a Laplace approximation on LoRA weights, a small set of low-rank adapter weights (Yang et al., 2024) (a workshop paper), and BNRM adds non-negative factor analysis to the Bradley-Terry model (Duan et al., 2026) (ICML 2026); RewardUQ found that model size and initialization matter more than the choice of uncertainty method. Human comparison noise is aleatoric, while acquisition by disagreement needs the epistemic part.
Modeling the oracle's bias. NAOD puts the judge's bias into acquisition (Section 35.2.4); LILO, which uses a language-model oracle, does not.
35.6.2 Population priors from pluralistic reward models #
Pluralistic and personalized reward models could serve as population priors for a new user of PBO. PAL (Chen et al., 2025c) (ICLR 2025) uses an ideal-point model, in which each user and option is a point in a shared latent space and preference falls with distance, with mixture modeling, and generalizes to new users from a few samples. The variational preference learning of Poddar et al. (2024) (NeurIPS 2024) infers a latent vector per user and suffers posterior collapse (the per-user latent stops carrying information) when each user has little data (Kim and Kim, 2026) (ICLR 2026). LoRe (Bose et al., 2025) (COLM 2025) and Cai et al. (2026) (SIGIR 2026) represent each user's reward as a weighted combination of shared basis functions, and on a benchmark of personalized reward models the best reached 75.94% accuracy (Ma et al., 2026) (COLM 2026). None chooses queries to learn a user's weights, so active per-user elicitation on a meta-learned low-rank prior is the natural next step for PBO to take from them (inference).
ActiveUltraFeedback's result also points to a distinction PBO has not yet drawn explicitly: the pairs that help locate the optimum are not the pairs that help learn a reusable utility model (inference). PBO work that builds population priors from earlier users is in the second situation; HOMI, for example, pretrains a prior from user models and was significantly better only at the second and third iterations (Liao et al., 2026) (CHI 2026).
Sources cited in Section 35.6 23
- Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
- Yang et al. (2026b) RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
- Ahmed and Ghasemi (2026) Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare
- Meta Platforms, Inc. (2026l) CHANGELOG (versions 1.2 to 1.3)
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Katkuri et al. (2026) Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning
- Wang et al. (2024b) RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
- Liu et al. (2025a) LiPO: Listwise Preference Optimization through Learning-to-Rank
- Liu et al. (2026a) Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph
- Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
- Rajagopalan et al. (2026) Personalized Image Generation via Human-in-the-loop Bayesian Optimization
- Lou et al. (2024) Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
- Coste et al. (2024) Reward Model Ensembles Help Mitigate Overoptimization
- Yang et al. (2024) Bayesian Reward Models for LLM Alignment
- Duan et al. (2026) Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
- Chen et al. (2025c) PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences
- Poddar et al. (2024) Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
- Kim and Kim (2026) Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
- Bose et al. (2025) LoRe: Personalizing LLMs via Low-Rank Reward Modeling
- Cai et al. (2026) One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
- Ma et al. (2026) Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models
35.7 Settled, contested, missing #
Settled. PBO and preference-based alignment share pairwise random utility likelihoods (Bradley-Terry, Thurstone, Plackett-Luce) and the contextual dueling-bandit framing, and acquisition rules move between them almost unchanged. The optimum of the KL-regularized objective is the reference policy tilted by , and DPO follows because the normalizer cancels in Bradley-Terry (Rafailov et al., 2023). The default PBO implementation uses the probit link, DPO the logistic one. Language models used as simulated users and judges are close to people in aggregate but unreliable for individuals, and biased by position, length, and self-preference. Language models acting directly as dueling-bandit optimizers trail classical algorithms in strong regret (Xia et al., 2025), and in scalar BO, agents performed no differently when their observations were replaced by random labels (Gupta et al., 2025). Only two of the eight systems in Table 35.1 are PBO in the strict sense.
Contested. Whether active selection of comparisons beats random selection for alignment: gains of 1% to 6% in win rate, 33% to 68% fewer labels, and an order of magnitude in simulated settings, against negligible gains in online DPO (Oh et al., 2026b); the reconciliation by the purpose of the query is our inference. Whether the exponential distortion of RLHF and DPO belongs to the algorithms (Gölz et al., 2025) or to a mismatch between data and reference policy (Oko et al., 2026). Whether the large label savings of Asghari et al. hold, given ablations that do not separate their two components and no replication. Whether the Bradley-Terry form is needed at all (Sun et al., 2025).
Missing. A PBO loop run with human comparisons and with language-model comparisons, reporting regret or sample efficiency for both. A measurement of how a judge's position bias propagates into acquisition decisions. A comparison, with human labels and a common budget, of double Thompson sampling, information-directed selection, quality-gap selection, and random selection for both reward-model RLHF and DPO. An evaluation with people of any preference-based prompt optimizer for text. A Gaussian process or kernel surrogate as the reward model of an RLHF loop at language-model scale. PBO methods built from alignment components (Plackett-Luce rankings, the aleatoric and epistemic split, oracle-bias models, pluralistic priors with active per-user elicitation), and an acquisition function that distinguishes pairs for finding the optimum from pairs for learning a reusable utility.
Sources cited in Section 35.7 7
- Rafailov et al. (2023) Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Xia et al. (2025) Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
- Gupta et al. (2025) LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
- Oh et al. (2026b) Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
- Gölz et al. (2025) Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
- Oko et al. (2026) Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner
- Sun et al. (2025) Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
Further reading #
- Rafailov et al. (2023) derive DPO from the KL-regularized objective in a few pages; the derivation in Section 35.1.3 follows theirs.
- Ouyang et al. (2022) describe the RLHF pipeline that made comparisons central to language models, with the details of how the comparisons were collected.
- Dwaracherla et al. (2024) and Oh et al. (2026b) are the two sides of active preference collection for alignment; read them together.
- Melikidze et al. (2026) is the most systematic comparison of acquisition rules for alignment data, and the clearest evidence that the purpose of a query changes which rule wins.
- Kobalczyk et al. (2026) is the cleanest example of a language model working only at the interface of a Gaussian process preference loop.
- Kirk et al. (2026), a preprint, measures in one experiment how a simulated user can be right about the population and wrong about each person.
- Siththaranjan et al. (2024) and Gölz et al. (2025) explain what a single reward does when many people answer.
References
- (2024). Online Bandit Learning with Offline Preference Data for Improved RLHF. arXiv (not accepted at TMLR). preprint Cited in §35.3
- (2026). Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare. Transactions on Machine Learning Research. Cited in §35.6
- (2026). Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences. arXiv. preprint Cited in §35.4
- (2026). Efficient Exploration at Scale. arXiv. preprint Cited in §35.3
- (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Cited in §35.2
- (2024b). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. arXiv. preprint Cited in §35.5
- (2024). A General Theoretical Paradigm to Understand Learning from Human Preferences. AISTATS 2024. Cited in §35.1 §35.4
- (2025). Online Preference Alignment for Language Models via Count-based Exploration. ICLR 2025. Cited in §35.3
- (2025). LoRe: Personalizing LLMs via Low-Rank Reward Modeling. Conference on Language Modeling (COLM 2025). Cited in §35.6
- (2026). One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment. SIGIR 2026. Cited in §35.6
- (2025). Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF. ICLR 2025. Cited in §35.3
- (2026a). Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference. IFAC Journal of Systems and Control. doi:10.1016/j.ifacsc.2026.100398. Cited in §35.3
- (2026). Multiple latent orderings better predict language model preferences. arXiv. preprint Cited in §35.2
- (2024). Preference Learning Algorithms Do Not Learn Preference Rankings. Advances in Neural Information Processing Systems. doi:10.52202/079017-3234. Cited in §35.4
- (2025a). Avoiding scaling in RLHF through Preference-based Exploration. Advances in Neural Information Processing Systems 38 (NeurIPS 2025). doi:10.52202/085713-5485. Cited in §35.3
- (2025c). PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences. ICLR 2025. Cited in §35.6
- (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. ICML 2024. Cited in §35.3
- (2025). LLM Routing with Dueling Feedback. arXiv (not accepted at ICLR 2026). preprint Cited in §35.3
- (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Cited in §35.4
- (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems. Cited in §35.1
- (2024). Reward Model Ensembles Help Mitigate Overoptimization. International Conference on Learning Representations. Cited in §35.1 §35.6
- (2025). Active Preference Optimization for Sample Efficient RLHF. Machine Learning and Knowledge Discovery in Databases. Research Track. doi:10.1007/978-3-032-06096-9_6. Cited in §35.3
- (2026). Optimal Design for Active Preference Learning with Biased LLM Judges. arXiv. preprint Cited in §35.2
- (2026). Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling. ICML 2026. Cited in §35.6
- (2023). AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback. Advances in Neural Information Processing Systems. doi:10.52202/075280-1308. Cited in §35.2
- (2024). Efficient Exploration for LLMs. ICML 2024. Cited in §35.3
- (2026). Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent Space. International Conference on Learning Representations. Cited in §35.2
- (2025). Thompson Sampling in Online RLHF with General Function Approximation. arXiv. preprint Cited in §35.3
- (2026). Cost-Aware Best-LLM Identification using Dueling Feedback. Advances in Neural Information Processing Systems. Cited in §35.3
- (2025). Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? NeurIPS 2025. Cited in §35.4 §35.7
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §35.5
- (2025). LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Findings of the Association for Computational Linguistics: EMNLP 2025. Cited in §35.2 §35.7
- (2026). A Unifying Lens on Reward Uncertainty in RLHF. arXiv. preprint Cited in §35.4
- (2024). Bayesian Preference Elicitation with Language Models. arXiv. preprint Cited in §35.2
- (2024). Reinforcement Learning from Human Feedback with Active Queries. TMLR. Cited in §35.3
- (2026). Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. ICLR 2026 RSI Workshop. workshop paper Cited in §35.3
- (2026). Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning. arXiv. preprint Cited in §35.6
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §35.3
- (2026). Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback. ICLR 2026. Cited in §35.6
- (2026). PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users. arXiv. preprint Cited in §35.2
- (2025). Active Task Disambiguation with LLMs. ICLR 2025. Cited in §35.2
- (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Cited in §35.2
- (2022). RL with KL penalties is better viewed as Bayesian inference. Findings of the Association for Computational Linguistics: EMNLP 2022. doi:10.18653/v1/2022.findings-emnlp.77. Cited in §35.4
- (2024a). A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules? International Conference on Machine Learning. Cited in §35.2
- (2026). Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design? arXiv. preprint Cited in §35.2
- (2025). Active Learning for Direct Preference Optimization. arXiv. preprint Cited in §35.3
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §35.4
- (2024b). Feel-Good Thompson Sampling for Contextual Dueling Bandits. International Conference on Machine Learning. Cited in §35.3
- (2025b). Eliciting Human Preferences with Language Models. ICLR 2025. Cited in §35.2
- (2026d). General Exploratory Bonus for Optimistic Exploration in RLHF. International Conference on Learning Representations. Cited in §35.3
- (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. Cited in §35.2
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Cited in §35.6
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §35.2 §35.6
- (2024a). On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization. Findings of the Association for Computational Linguistics: EMNLP 2024. doi:10.18653/v1/2024.findings-emnlp.940. Cited in §35.4
- (2024b). Prompt Optimization with Human Feedback. ICML 2024 MHFAIA Workshop (no formal proceedings). workshop paper Cited in §35.3
- (2024a). Large Language Models to Enhance Bayesian Optimization. International Conference on Learning Representations. Cited in §35.2
- (2024b). Sample-Efficient Alignment for LLMs. NeurIPS 2024 LanGame Workshop. workshop paper Cited in §35.3
- (2025a). LiPO: Listwise Preference Optimization through Learning-to-Rank. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). doi:10.18653/v1/2025.naacl-long.121. Cited in §35.5 §35.6
- (2026a). Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph. arXiv. preprint Cited in §35.6
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §35.6
- (2024). Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown. arXiv (withdrawn from ICLR 2025). preprint Cited in §35.6
- (2026). Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization. COLM 2026. Cited in §35.6
- (2025). MAPLE: A Framework for Active Preference Learning Guided by Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v39i26.34964. Cited in §35.2
- (2025). Sample Efficient Preference Alignment in LLMs via Active Exploration. COLM 2025. Cited in §35.3
- (2026). ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning. ICML 2026. Cited in §35.3
- (2024). Deep Bayesian Active Learning for Preference Modeling in Large Language Models. NeurIPS 2024. Cited in §35.3
- (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Cited in §35.5
- (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Cited in §35.4
- (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Cited in §35.6
- (2024). Active Preference Learning for Large Language Models. ICML 2024. Cited in §35.2 §35.3
- (2024). Nash Learning from Human Feedback. ICML 2024. Cited in §35.4
- (2026). Efficient Exploration for Iterative Nash Preference Optimization. arXiv. preprint Cited in §35.3
- (2026). CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM. International Conference on Machine Learning (ICML 2026). Cited in §35.3
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Cited in §35.2
- (2026b). Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs. ICLR 2026 Workshop: I Can't Believe It's Not Better (ICBINB). workshop paper Cited in §35.3 §35.7
- (2026). Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner. International Conference on Machine Learning. Cited in §35.4 §35.7
- (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems. Cited in §35.1
- (2024). LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS 2024. Cited in §35.2
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Cited in §35.6
- (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Cited in §35.2
- (2023). Active Preference Inference using Language Models and Probabilistic Reasoning. NeurIPS 2023 FMDM Workshop. workshop paper Cited in §35.2
- (2024). Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning. Advances in Neural Information Processing Systems 37 (NeurIPS 2024). doi:10.52202/079017-1664. Cited in §35.6
- (2026). Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models. Nature Communications. doi:10.1038/s41467-025-67998-6. Cited in §35.2
- (2026). T-POP: Test-Time Personalization with Online Preference Feedback. International Conference on Machine Learning. Cited in §35.3
- (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. Cited in §35.1 §35.4 §35.7
- (2026). Personalized Image Generation via Human-in-the-loop Bayesian Optimization. International Conference on Machine Learning. Cited in §35.6
- (2026). Bayesian Optimization of Catalysis With In-Context Learning. ACS Central Science. doi:10.1021/acscentsci.5c02418. Cited in §35.2
- (2026). Large language models as uncertainty-calibrated optimizers for experimental discovery. Nature Machine Intelligence. doi:10.1038/s42256-026-01283-z. Cited in §35.2
- (2026). When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model. arXiv. preprint Cited in §35.2
- (2025). Reproducibility Study of Large Language Model Bayesian Optimization. arXiv. preprint Cited in §35.2
- (2026). Beyond expert users: agents should help users construct preferences, not just elicit them. Conference on Language Modeling (COLM 2026). Cited in §35.2
- (2024). Optimal Design for Reward Modeling in RLHF. arXiv. preprint Cited in §35.3
- (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Cited in §35.2
- (2026). Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2026.acl-long.2192. Cited in §35.2
- (2025a). Active Reward Modeling: Adaptive Preference Labeling for Large Language Model Alignment. International Conference on Machine Learning. Cited in §35.3
- (2025). Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics. Cited in §35.2
- (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Cited in §35.4
- (2024). The Importance of Online Data: Understanding Preference Fine-tuning via Coverage. NeurIPS 2024. Cited in §35.4
- (2025). Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives. ICLR 2025. Cited in §35.4 §35.7
- (2026). MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization. arXiv. preprint Cited in §35.3
- (2024). Generalized Preference Optimization: A Unified Approach to Offline Alignment. International Conference on Machine Learning. Cited in §35.1 §35.4
- (2026a). Calibrated Preference Learning: The Case of Label Ranking. International Conference on Machine Learning (ICML 2026). Cited in §35.5
- (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. Cited in §35.3
- (2023a). Is RLHF More Difficult than Standard RL? NeurIPS 2023. Cited in §35.4
- (2024a). Large Language Models are not Fair Evaluators. ACL 2024. Cited in §35.2
- (2024b). RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback. International Conference on Machine Learning. Cited in §35.6
- (2025). Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization. arXiv. preprint Cited in §35.4
- (2024). Making RL with Preference-based Feedback Efficient via Randomization. ICLR 2024. Cited in §35.3
- (2026). LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization. Findings of the Association for Computational Linguistics: ACL 2026. doi:10.18653/v1/2026.findings-acl.490. Cited in §35.3
- (2025). Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of the Association for Computational Linguistics: ACL 2025. doi:10.18653/v1/2025.findings-acl.519. Cited in §35.2 §35.7
- (2025). Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF. ICLR 2025. Cited in §35.3
- (2024). Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint. ICML 2024. Cited in §35.4
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §35.5
- (2025a). Investigating Non-Transitivity in LLM-as-a-Judge. International Conference on Machine Learning. Cited in §35.2
- (2024). Bayesian Reward Models for LLM Alignment. ICLR 2024 SeT LLM Workshop; ICML 2024 SPIGM Workshop. workshop paper Cited in §35.6
- (2026a). Quantifying and Mitigating Self-Preference Bias of LLM Judges. arXiv. preprint Cited in §35.2
- (2026b). RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models. EurIPS 2025 EIML Workshop. workshop paper Cited in §35.6
- (2026). Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery. ICLR 2026. Cited in §35.2
- (2024a). Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought. International Conference on Machine Learning. Cited in §35.3
- (2024b). Self-Exploring Language Models: Active Preference Elicitation for Online Alignment. TMLR. Cited in §35.3
- (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §35.6
- (2025). Sharp Analysis for KL-Regularized Contextual Bandits and RLHF. NeurIPS 2025. Cited in §35.4
- (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems. doi:10.52202/075280-2020. Cited in §35.2
- (2023). Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. International Conference on Machine Learning. Cited in §35.4
- (2024). How Reliable is Your Simulator? Analysis on the Limitations of Current LLM-based User Simulators for Conversational Recommendation. Companion Proceedings of the ACM Web Conference 2024. doi:10.1145/3589335.3651955. workshop paper Cited in §35.2