Bayesian Optimization
Part VIII: Neighbors in Computing
中文

Preferences and Large Language Models

The most visible use of pairwise comparisons in computing today is not preferential Bayesian optimization (PBO). It is the alignment of large language models, where people (and increasingly other models) read two responses to the same prompt and say which is better, millions of times, and the answers are used to tune the model. The question asked of the annotator is the question Chapter 19 asks of a person tuning a design: which of these two do you prefer? The likelihood used to learn from the answer is, in most alignment work, the Bradley-Terry model of Section 16.4.

From 2023 to September 2026 the two fields met in two directions. Language models moved into preference loops, as priors, simulated users, translators from text to preference labels, feature extractors, and generators of candidates. In the other direction, tools from dueling bandits and Bayesian active learning were applied to prompt optimization, to choosing among outputs and models, and to choosing which comparisons to collect for training. For this book's argument, alignment is where the measurement question was asked at scale: how far a model judge departs from a person, and which comparisons are worth asking for (Section 45.1). The reported gains range from 1% to 6% in win rate to roughly an order of magnitude in labels saved, but most results replace people with a reward-model simulator or a language-model judge, and some careful studies find that choosing comparisons cleverly barely beats choosing them at random. The chapter explains how alignment learns from comparisons, follows the two directions, separates the likelihood the fields share from seven analogies that break, and ends with what has flowed back.

35.1 How language models learn from comparisons #

A reader who knows Gaussian process preference learning already knows most of the machinery. This section introduces the vocabulary of alignment in that reader's terms.

35.1.1 Policies, rewards, and the Bradley-Terry reward model #

A language model, given a prompt xx, produces a response yy by sampling words one at a time, which defines a probability distribution over responses, π(y ∣ x)\pi(y \given x). Alignment research borrows the word policy from reinforcement learning for this distribution: a rule for choosing actions, where the action is the entire response. The model before preference tuning, usually already fine-tuned on written demonstrations, is the reference policy πref\pi_{\text{ref}}.

A reward model is a separate network that maps a prompt and a response to a single number r(x,y)r(x, y), often a copy of the language model with its output layer replaced by one scalar (Ouyang et al., 2022). It plays exactly the role of the latent utility ff in Chapter 18: nobody observes it, and it is learned only from comparisons. If an annotator is shown responses yy and y′y' to prompt xx, the Bradley-Terry model says

P(y≻y′ ∣ x)=sigmoid⁡(r(x,y)−r(x,y′)),sigmoid⁡(z)=11+e−z,\Prob(y \succ y' \given x) = \operatorname{sigmoid}\big(r(x, y) - r(x, y')\big), \qquad \operatorname{sigmoid}(z) = \frac{1}{1 + e^{-z}},
(35.1)

where sigmoid⁡\operatorname{sigmoid} is the logistic function (language-model papers write it σ\sigma, a letter this book reserves for standard deviations). A reward model is trained by maximizing the log-likelihood of comparisons, each a prompt with a preferred response ywy_w and a rejected one yly_l:

L(r)=−∑(x,yw,yl)log⁡sigmoid⁡(r(x,yw)−r(x,yl)).\mathcal{L}(r) = -\sum_{(x, y_w, y_l)} \log \operatorname{sigmoid}\big(r(x, y_w) - r(x, y_l)\big).
(35.2)

Set beside Equation (19.1), the only differences are the link (logistic rather than probit, two noise distributions of the same random utility family, Section 16.5), the absence of a prior on rr, and the scale: a large neural network fitted to tens of thousands to millions of comparisons instead of a Gaussian process fitted to a few dozen. The ranking extension, in which the probability of an ordering of KK responses is a product of softmax choices from the remaining items, is the Plackett-Luce model.

35.1.2 RLHF and the KL-regularized objective #

Reinforcement learning from human feedback (RLHF) uses such a reward model to tune a policy. Christiano et al. (2017) introduced it for games and simulated robots, and Ouyang et al. (2022) applied it to instruction-following language models. It samples responses and collects human comparisons, fits a reward model with Equation (35.2), and then fine-tunes the policy to maximize, for each prompt xx and on average over prompts,

Ey∼π[r(x,y)]−β KL⁡(π ∥ πref),\E_{y \sim \pi}\big[r(x, y)\big] - \beta\,\KL\big(\pi \,\|\, \pi_{\text{ref}}\big),
(35.3)

where π\pi and πref\pi_{\text{ref}} are short for π(⋅ ∣ x)\pi(\cdot \given x) and πref(⋅ ∣ x)\pi_{\text{ref}}(\cdot \given x). The second term is the KL divergence of Section 6.2, how far the tuned policy has moved from the reference, and β>0\beta > 0 sets how much movement the reward must pay for. Without it, the policy would drift toward whatever text the reward model overrates, a failure called reward over-optimization: the learned reward keeps rising while true quality falls (Coste et al., 2024). Notice what Equation (35.3) asks for: not the single best response, but a distribution over responses that scores well on average while staying close to the reference, the first analogy that breaks in Section 35.4.2.

35.1.3 DPO and the implicit reward #

Direct preference optimization (DPO) removes the separate reward model and the reinforcement learning step. Rafailov et al. (2023) observed that Equation (35.3) has a closed-form optimum, and that the Bradley-Terry likelihood can then be written directly in terms of the policy.

Derivation From the KL-regularized objective to the DPO loss

Fix one prompt xx and drop it from the notation. Write π(y)\pi(y) for the policy and sums over all responses yy.

  1. The objective of Equation (35.3) for this prompt is J(π)=∑yπ(y) r(y)−β∑yπ(y)log⁡π(y)πref(y)J(\pi) = \sum_y \pi(y)\, r(y) - \beta \sum_y \pi(y) \log \frac{\pi(y)}{\pi_{\text{ref}}(y)}, by the definition of the expectation and of the KL divergence.
  2. Define Z=∑yπref(y) er(y)/βZ = \sum_y \pi_{\text{ref}}(y)\, e^{r(y)/\beta} and the distribution π⋆(y)=πref(y) er(y)/β/Z\pi^\star(y) = \pi_{\text{ref}}(y)\, e^{r(y)/\beta} / Z. Factoring −β-\beta out of both terms of step 1 gives J(π)=−β∑yπ(y)log⁡π(y)πref(y) er(y)/βJ(\pi) = -\beta \sum_y \pi(y) \log \frac{\pi(y)}{\pi_{\text{ref}}(y)\, e^{r(y)/\beta}}. Writing πref(y) er(y)/β=Z π⋆(y)\pi_{\text{ref}}(y)\, e^{r(y)/\beta} = Z\, \pi^\star(y) inside the logarithm splits it into log⁡π(y)π⋆(y)−log⁡Z\log \frac{\pi(y)}{\pi^\star(y)} - \log Z, and because ∑yπ(y)=1\sum_y \pi(y) = 1 the constant contributes −log⁡Z-\log Z once: J(π)=−β KL⁡(π ∥ π⋆)+βlog⁡ZJ(\pi) = -\beta\, \KL(\pi \,\|\, \pi^\star) + \beta \log Z.
  3. A KL divergence is never negative and is zero only when the two distributions are equal, so JJ is maximized by π=π⋆\pi = \pi^\star:
    π⋆(y)=1Z πref(y) er(y)/β.\pi^\star(y) = \frac{1}{Z}\, \pi_{\text{ref}}(y)\, e^{r(y)/\beta}.
    (35.4)
  4. Take logarithms of Equation (35.4) and solve for the reward:
    r(y)=βlog⁡π⋆(y)πref(y)+βlog⁡Z.r(y) = \beta \log \frac{\pi^\star(y)}{\pi_{\text{ref}}(y)} + \beta \log Z.
    (35.5)
  5. Restore the prompt and substitute Equation (35.5) into the Bradley-Terry model Equation (35.1). The difference r(x,yw)−r(x,yl)r(x, y_w) - r(x, y_l) contains βlog⁡Z(x)\beta \log Z(x) twice with opposite signs, because ZZ depends on the prompt but not on the response, so it cancels. The intractable sum over all possible responses disappears.
  6. What remains is a likelihood in which the policy itself plays the role of the reward. Write r^π(x,y)=βlog⁡π(y ∣ x)−βlog⁡πref(y ∣ x)\hat r_\pi(x, y) = \beta \log \pi(y \given x) - \beta \log \pi_{\text{ref}}(y \given x) for this policy-defined reward. Maximizing the Bradley-Terry likelihood over a data set of comparisons is then the DPO loss:
    LDPO(π)=−∑(x,yw,yl)log⁡sigmoid⁡(r^π(x,yw)−r^π(x,yl)).\mathcal{L}_{\text{DPO}}(\pi) = -\sum_{(x, y_w, y_l)} \log \operatorname{sigmoid}\big(\hat r_\pi(x, y_w) - \hat r_\pi(x, y_l)\big).
    (35.6)

In the words of its title, the policy "is secretly a reward model": the quantity r^π(x,y)\hat r_\pi(x, y) is called the implicit reward, and Equation (35.6) is Equation (35.2) with r^π\hat r_\pi in place of rr, with a Plackett-Luce version for rankings in the paper. Two relatives recur below, both members of one family of convex losses (Tang et al., 2024): IPO (Azar et al., 2024) uses a squared loss, and SLiC a hinge loss. And Equation (35.4) says that the optimal policy is the reference policy reweighted by er/βe^{r/\beta}, a tilt of what the model already does rather than a jump to the best response.

35.1.4 Where comparisons come from #

Offline preference learning uses a fixed data set gathered in advance; online learning samples new responses from the current policy during training and sends them for labeling; active learning also chooses which prompts and pairs are worth labeling. That choice is an acquisition function in the sense of Chapter 12, and it is where tools from PBO and dueling bandits enter alignment. A dueling bandit (Section 21.1) chooses two arms per round and observes which one wins; a contextual dueling bandit first observes a context, here the prompt. Much of the evidence below uses an LLM judge, a language model prompted to compare two responses in place of a person, and reports a win rate, the fraction of head-to-head comparisons against a fixed baseline model that the tuned model wins.

Key idea Same question, different job

Alignment and PBO both learn a latent score from answers to "which is better?" with a pairwise random utility likelihood. Alignment then uses the score to reshape a distribution over text near a reference model; PBO uses it to find one best design. Most of what follows comes back to that difference.

Sources cited in Section 35.1 6
  1. Ouyang et al. (2022) Training language models to follow instructions with human feedback
  2. Christiano et al. (2017) Deep reinforcement learning from human preferences
  3. Coste et al. (2024) Reward Model Ensembles Help Mitigate Overoptimization
  4. Rafailov et al. (2023) Direct Preference Optimization: Your Language Model is Secretly a Reward Model
  5. Tang et al. (2024) Generalized Preference Optimization: A Unified Approach to Offline Alignment
  6. Azar et al. (2024) A General Theoretical Paradigm to Understand Learning from Human Preferences

35.2 Language models inside the preference loop #

Between 2023 and 2026 language models joined preference loops in five jobs: a prior or warm start, a simulated user, a translator from text to preference labels, a feature extractor, and a generator of candidates. The systems that worked best kept a classical probabilistic model in charge of uncertainty and of choosing queries, and let the language model work only at the interface; systems in which the language model itself acted as the optimizer fell behind classical dueling-bandit algorithms in strong regret.

35.2.1 Translators: from words to preference signals #

PEBOL (Austin et al., 2024a) elicits preferences in conversational recommendation with an independent Beta-distributed utility per item, the conjugate model of Section 5.2: a natural language inference model turns the user's "yes" or "no" about an aspect into a likelihood, and Thompson sampling or an upper confidence bound (Section 12.5, Section 12.4) decides what GPT-3.5 should ask next. After 10 turns, its mean reciprocal rank at 10 (the average of one over the rank of the user's target item in the top-10 list) was 0.27 on Yelp against 0.12 for a monolithic GPT-3.5 elicitor, 0.18 against 0.09 on MovieLens, and 0.17 against 0.11 on Recipe-MPR; the best monolithic baseline, Gemini-Pro, reached 0.17. The authors attribute the monolithic models' failure to over-exploitation, in severe cases asking the same question again. The users were 100 GPT-3.5 simulations per data set, told in advance which item they liked.

OPEN (Handa et al., 2024), a 2024 preprint, has a language model extract and rank features of a domain to initialize the prior of a linear Bradley-Terry utility, chooses pairwise queries by the expected information gain of Section 6.4 over a particle-filter posterior (a cloud of weighted samples), and has the language model rewrite each abstract comparison as a natural question. With people on the Prolific platform, recommending New York Times articles, it beat elicitation by the language model alone and by experimental design alone; without the rewriting, users found it markedly harder to express their preferences accurately, and mental demand was similar across methods. MAPLE (Mahmud et al., 2025) (AAAI 2025) turns natural-language feedback into samples or ranges for the linear weights of abstract concepts, with a Bradley-Terry likelihood over pairwise rankings of trajectories and Markov chain Monte Carlo (Section 17.5), evaluated with modeled humans.

LILO (Kobalczyk et al., 2026) (ICML 2026) is the closest to PBO in the strict sense. A language model translates free-text feedback and textual priors into pairwise labels for a probit pairwise Gaussian process over a composite utility u(x)=g(f(x))u(\vx) = g(f(\vx)), where ff gives an experiment's measured outcomes and gg the decision maker's utility over them, the setting of preference exploration (Lin et al., 2022), and qEUBO, the expected utility of the best option for queries of qq options (Section 19.4), selects which outcomes to have compared. Across 10 environments it beat pure language-model optimizers and PBO baselines, and in some settings one message from the decision maker matched or exceeded 8 to 16 pairwise comparisons; the problems had 5 to 8 dimensions, and the decision maker was simulated by Llama-3.3-70B. Its precursors include preregistered experiments with people in which questions generated by a language model elicited answers often more informative than prompts or labels users wrote themselves, with less reported effort (Li et al., 2025b); a NeurIPS 2023 workshop paper that chose questions by expected entropy reduction (Piriyakulkij et al., 2023); and, by LILO's first author, Bayesian experimental design over candidate solutions sampled from a language model (Kobalczyk et al., 2025) (ICLR 2025).

35.2.2 Priors, features, candidates, and the optimizer itself #

Priors and warm starts. Besides OPEN and MAPLE, Eichelbeck et al. (2026) (ICLR 2026) run PBO in an autoencoder's latent space with a prior a language model generates from interviews, with simulated users only. On the scalar side (Section 30.6), LLAMBO (Liu et al., 2024a) (ICLR 2024) uses a language model for warm starts, as a surrogate, and to sample candidates; a 2025 reproduction preprint found that warm starting "substantially improves early regret behaviour and reduces variance across runs" but that the surrogate is "weaker than GP or SMAC as a pure single task regressor" (SMAC uses a random forest), and its abstract attributes LLAMBO to "Daxberger et al. (2024)", which does not match the original authors (Rychert et al., 2025). LGBO (Yuan et al., 2026) (ICLR 2026) adds a language model's preferences about regions to the surrogate mean, with a worst-case guarantee, faster convergence when the preferences agree with the objective, and one wet-lab experiment, for a scalar objective. The evidence against is pointed: language-model agents in scalar BO performed no differently when their observed outcomes were replaced by randomly permuted labels, while linear bandits and Gaussian process optimization consistently won (Gupta et al., 2025) (EMNLP 2025 Findings), and the better starting point a language-model adviser suggested was a default configuration (Rodrigues et al., 2026), a 2026 preprint.

Feature extractors and surrogates. Language models help BO over molecules "only if they have been pretrained or finetuned with domain-specific" data (Kristiadi et al., 2024a) (ICML 2024); BO-ICL (Ramos et al., 2026), first posted in 2023 and published in 2026, found a near-optimal catalyst among 3,700 candidates within 6 iterations by in-context regression; and Ranković et al. (2026) (Nature Machine Intelligence 2026) train embeddings and a Gaussian process jointly, matching conventional BO with a median of 41% fewer iterations across 23 chemistry and materials tasks. All three use scalar feedback; with preference feedback, the instances of "language model embedding plus a probabilistic head" are APOHF and the method of Dwaracherla et al. (Section 35.3), neither with a Gaussian process.

Candidates and the optimizer itself. APPO (Li et al., 2026f) (CHI 2026) has a language model rewrite text-to-image prompts from a user's binary preferences, and PDO and Duel-Evolve (Section 35.3) let one language model both generate candidates and judge them. Xia et al. (2025) (ACL 2025 Findings) asked top language models to act as dueling-bandit algorithms in context. They quickly brought the best arm into the duels, giving low short-term weak regret (which counts the better of the two arms), but in strong regret (which counts both) "an optimality gap still exists" with classical algorithms; they "struggle to converge and consistently exploit even when explicitly prompted to do so". A hybrid, LEAD, wraps a classical algorithm around the language model and inherits its guarantees.

Table 35.1 sorts the main systems by whether they are PBO in the strict sense used in this book: a probabilistic surrogate that learns a latent utility over a continuous or structured design space from pairwise or ranked feedback, with an acquisition function choosing the queries.

Table 35.1 Systems that put a language model in a preference loop, and whether each is PBO in the strict sense.
System Role of the language model Probabilistic model and query rule Feedback Strict PBO? Who answered in the evaluation
PEBOL (RecSys 2024) asks questions; an inference model turns answers into a likelihood independent Beta-Bernoulli item utilities; Thompson sampling or UCB yes or no no: independent-arm bandit 100 GPT-3.5 simulated users per data set
OPEN (2024 preprint) extracts features, initializes the prior, rewrites questions linear Bradley-Terry utility; particle filter; expected information gain pairwise partly: no GP people on Prolific
MAPLE (AAAI 2025) supplies weight priors, interprets language feedback linear weights over concepts; Bradley-Terry; MCMC pairwise rankings plus language partly: as for OPEN modeled humans
LILO (ICML 2026) translates free text into pairwise labels probit pairwise GP on a composite utility; qEUBO language turned into pairwise labels yes Llama-3.3-70B simulated decision maker, 5 to 8 dimensions
Eichelbeck et al. (ICLR 2026) generates the prior from interviews PBO in an autoencoder latent space pairwise yes simulated users only
LGBO (ICLR 2026) states regional preferences shift of the GP mean scalar plus language-model preferences no: scalar BO benchmarks and one wet-lab experiment
APOHF (workshop paper) frozen embeddings as features neural network, Bradley-Terry; greedy plus UCB pairwise no: neural dueling bandit simulated (validation accuracy, image similarity)
Duel-Evolve (workshop paper) generates candidates and judges them Bayesian Bradley-Terry; double Thompson sampling the model's own pairwise judgments no: test-time search ground-truth benchmark accuracy

Only LILO and the method of Eichelbeck et al. meet the strict definition, so calling all of these "PBO with language models" would overstate how much of this intersection Gaussian process preference models occupy.

35.2.3 Language for goals, comparisons for judgments #

If a language model can carry a person's words into the loop, should the person still be asked to compare? In the design studies of Chapter 32, natural language and explicit constraints gave the same optimization performance, with lower workload for language and a stronger sense of agency for constraints, and 90.9% of 187 natural-language requests described a desired outcome rather than a parameter value (Niwa et al., 2025). In a preprint, Peng et al. (2026) found that 20 people judging the same 600 pairs of generated interfaces agreed with one another at a Krippendorff's α\alpha of only 0.25 (two of them made the same choice on 62.4% of pairs on average), yet for 12 new users, personalization from just 8 pairwise judgments beat every baseline, including the users' own written preferences, with an aggregate win rate of 60.35% (Section 32.4).

The model side points the same way. Language models inferring a user's preferences over several turns fall "far short" of normative Bayesian updating, and training them to imitate a Bayesian model improves this and generalizes (Qiu et al., 2026) (Nature Communications, January 2026); and on a simulated-user benchmark for helping users construct preferences, no frontier model exceeded 56% accuracy within 5 turns (Saracay et al., 2026) (COLM 2026). Together these results support a division of labor: language for goals, constraints, and priors; comparisons for fine judgments; and posterior updating left to an explicit probabilistic model (inference), as in the case for comparisons over ratings in Section 16.1.

35.2.4 Language models as simulated users and judges #

Many systems in this chapter were evaluated with a language model standing in for the person. The evidence from 2023 to 2026 points one way: close to people in aggregate, unreliable for individuals, and biased in systematic directions. In Table 35.2, Spearman correlation and Kendall's τ\tau both measure how similarly two lists are ordered (1 for identical order, 0 for no relation), and agreement rates count how often two judges pick the same option.

Table 35.2 How faithful language models are as judges and simulated users, in aggregate and for individuals.
Study Setting In aggregate For individuals, or biases
Zheng et al., NeurIPS 2023 Datasets and Benchmarks GPT-4 judging MT-Bench and Chatbot Arena agreement with people "over 80%", the same as between people position bias, verbosity bias, self-enhancement bias
Dubois et al., NeurIPS 2023, AlpacaFarm GPT-4 annotator with a single prompt 65% agreement with people against 66% between people; method rankings correlate with those from human data at Spearman 0.98, at about 50 times lower cost does not reproduce the variability of human labels or reward over-optimization; reproducing it needed a pool of simulated annotators and labels flipped at random with probability 0.25
Wang et al., ACL 2024 ChatGPT as evaluator not reported after swapping the order of responses, Vicuna-13B beat ChatGPT on 66 of 80 queries
Panickssery, Bowman, and Feng, NeurIPS 2024 fine-tuning changes self-recognition not reported self-recognition ability correlates linearly with the strength of self-preference
Muldrew et al., ICML 2024 50 prompts labeled twice GPT-4 agrees with itself more than 90% of the time GPT-3.5-turbo about 60%
Kirk et al., 2026 preprint, PRISM-X 530 people, each ranking four models; GPT-4o as simulated user, either ranking the person's transcripts or also holding the conversations model scores fitted to simulated and to human rankings correlate at r = 0.99 and 0.98 mean Kendall's τ with the person's own ranking of 0.22 (ranking only) and 0.11 (conversing), against 0.57 for people's own consistency; strong position effects (Figure 35.1); more sycophantic than people
Kuric, Demcak, and Krajcovic, 2026 preprint 29 real design preference tests (n = 2073), 78 tasks top-choice agreement 53%; GPT 5.2 reaches 65% with no significant gain in fidelity simulated and real distributions differ significantly in 44% of tasks; single-persona simulation deviates significantly in 91%; temperature and top-p have no significant effect
Xu et al., ICML 2025 judges on AlpacaEval round-robin plus Bradley-Terry raises Spearman correlation with Chatbot Arena from 95.0% to 96.4%, Kendall from 82.1% to 86.3% judgments are intransitive; rankings depend on the chosen baseline model
Shi et al., IJCNLP-AACL 2025 15 judges, over 150,000 instances not reported position bias correlates weakly with prompt length and is "strongly affected by the quality gap between solutions"
Yang et al., 2026 preprint 20 language models, pairs of responses of equal quality not reported capability and low self-preference bias are "often uncorrelated, or even negatively correlated"; structured multi-dimensional evaluation cuts the bias by 31.5% on average
Chawla, Thompson, and Young, 2026 preprint 7 models, 4 tasks not reported intransitivity cannot be explained by a single ordering under any monotone link; a mixture of several latent orderings fits better
Seshadri et al., 2026 preprint simulated users in agent evaluation not reported agent success rates differ by up to 9 percentage points between simulators; fidelity is worse for some dialect groups
Zhu, Huang, and Sang, WWW 2024 workshop simulated users in conversational recommendation not reported data leakage in the dialogue history and the simulator's replies inflates results; PEBOL relies on exactly this kind of evaluation

The sources are, in order, Zheng et al. (2023), Dubois et al. (2023), Wang et al. (2024a), Panickssery et al. (2024), Muldrew et al. (2024), Kirk et al. (2026), Kuric et al. (2026), Xu et al. (2025a), Shi et al. (2025), Yang et al. (2026a), Chawla et al. (2026), Seshadri et al. (2026), and Zhu et al. (2024). Position bias is a preference for whichever response is shown first (or second); verbosity bias a preference for longer responses; self-enhancement or self-preference bias a judge's preference for text it generated itself; sycophancy a tendency to tell the user what they appear to want to hear. Figure 35.1 draws the clearest case, PRISM-X.

Agreement with people00.250.500.751model scores, simulated vs human (r)0.98 to 0.99each simulated ranking vs the person's own (τ)0.22people's ratings vs their own rankings (τ)0.57Ranked best, by position shown0%10%20%30%40%50%24.133.8shown 1st28.615.0shown 4thpeopleranks transcriptsholds conversations25%, no position effect
Agreement with people00.250.500.751model scores, simulated vs human (r)0.98 to 0.99each simulated ranking vs the person's own (τ)0.22people's ratings vs their own rankings (τ)0.57Ranked best, by position shown0%10%20%30%40%50%24.133.8shown 1st28.615.0shown 4thpeopleranks transcriptsholds conversations25%, no position effect
Figure 35.1 Right about the population, wrong about the person, in PRISM-X (Kirk et al., a 2026 preprint): 530 people each ranked four models, and GPT-4o, as a simulated user, ranked the same models either from the person's transcripts only or after holding the conversations itself. Agreement with people: model scores fitted to simulated rankings correlate with those fitted to human rankings at r = 0.98 to 0.99 (the two conditions), while each simulated ranking agrees with that person's own at a mean Kendall's τ of 0.22 or 0.11, against 0.57 for the consistency of people's own ratings and rankings. Ranked best: the share of trials in which the model shown first, or fourth, of four was ranked best; with no position effect each would be 25%. Choose a condition to highlight it. All numbers as reported (Kirk et al., 2026).

Things to try:

  • With Simulated user set to Ranks transcripts only, compare the top bar of Agreement with people with the bar below it: near-perfect agreement about which model is better overall, and a per-person agreement of 0.22.
  • Switch to Also holds the conversations. Per-person agreement halves to 0.11, and the share of trials in which the first-shown model wins rises from 33.8% to 44.9%, against 24.1% for people.

For PBO that uses a language model as its preference oracle, this evidence means the observation model is misspecified (inference). Language-model labels are faithful in aggregate and wrong for individuals (PRISM-X, Kuric et al.), too consistent (AlpacaFarm needed 25% random flips to reproduce human-like over-optimization), dependent on presentation order (Wang et al.), and biased toward the judge's own text (Panickssery et al.). A probit or Bradley-Terry likelihood treats every deviation as independent noise with a fixed variance, but most of these deviations are bias, and a position bias that varies with the quality gap (Shi et al.) makes the noise variance depend on the utility difference, which a fixed noise scale cannot express (Section 27.2 lists heteroscedastic models). The most direct remedy is the balanced position calibration of Wang et al.: average the judgments from both orders. NAOD (Du et al., 2026), a preprint posted on September 30, 2026, is the first method we found that puts judge bias into the design of queries: arguing that active acquisition exposes judge bias that remains after calibration, it lowered average proxy policy regret by 29.1% on Chatbot Arena data with 17 judges, against a matched design that targets information alone, and showed that representation error "can reverse an oracle design advantage". LILO does not model judge bias.

As of September 2026 we found no study that runs the same PBO loop with human comparisons and with language-model comparisons and reports regret or sample efficiency for both, and none that measures how position bias propagates into acquisition decisions (for example, whether averaging over both orders changes the sequence of selected queries). Until such a study exists, a PBO result obtained with a language-model oracle is a result about that oracle; the minimal check is to rerun a few sessions with people and with both presentation orders, and to report how the selected queries change (inference). The gap is not peculiar to language models: in a retinal-implant study within PBO, only about 50% of human choices agreed with a simulated agent that was not a language model, yet 16 of 17 participants preferred the optimized result (Schoinas et al., 2025) (Section 33.5, Section 31.7).

35.2.5 The pattern #

From PEBOL in 2023 to LILO in 2026, the model of uncertainty stayed classical and small (PEBOL's Beta-Bernoulli model, the linear Bradley-Terry models of OPEN and MAPLE, LILO's pairwise Gaussian process), and the language model worked only at the interfaces: features, prior initialization, wording of questions, and translation of labels. Every approach that handed query selection to the language model fell behind classical algorithms (inference).

Sources cited in Section 35.2 38
  1. Austin et al. (2024a) Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation
  2. Handa et al. (2024) Bayesian Preference Elicitation with Language Models
  3. Mahmud et al. (2025) MAPLE: A Framework for Active Preference Learning Guided by Large Language Models
  4. Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback
  5. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  6. Li et al. (2025b) Eliciting Human Preferences with Language Models
  7. Piriyakulkij et al. (2023) Active Preference Inference using Language Models and Probabilistic Reasoning
  8. Kobalczyk et al. (2025) Active Task Disambiguation with LLMs
  9. Eichelbeck et al. (2026) Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent Space
  10. Liu et al. (2024a) Large Language Models to Enhance Bayesian Optimization
  11. Rychert et al. (2025) Reproducibility Study of Large Language Model Bayesian Optimization
  12. Yuan et al. (2026) Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery
  13. Gupta et al. (2025) LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
  14. Rodrigues et al. (2026) When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model
  15. Kristiadi et al. (2024a) A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
  16. Ramos et al. (2026) Bayesian Optimization of Catalysis With In-Context Learning
  17. Ranković et al. (2026) Large language models as uncertainty-calibrated optimizers for experimental discovery
  18. Li et al. (2026f) Preference-Guided Prompt Optimization for Text-to-Image Generation
  19. Xia et al. (2025) Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
  20. Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
  21. Peng et al. (2026) Efficient Personalization of Generative User Interfaces
  22. Qiu et al. (2026) Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
  23. Saracay et al. (2026) Beyond expert users: agents should help users construct preferences, not just elicit them
  24. Zheng et al. (2023) Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
  25. Dubois et al. (2023) AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
  26. Wang et al. (2024a) Large Language Models are not Fair Evaluators
  27. Panickssery et al. (2024) LLM Evaluators Recognize and Favor Their Own Generations
  28. Muldrew et al. (2024) Active Preference Learning for Large Language Models
  29. Kirk et al. (2026) PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
  30. Kuric et al. (2026) Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design?
  31. Xu et al. (2025a) Investigating Non-Transitivity in LLM-as-a-Judge
  32. Shi et al. (2025) Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
  33. Yang et al. (2026a) Quantifying and Mitigating Self-Preference Bias of LLM Judges
  34. Chawla et al. (2026) Multiple latent orderings better predict language model preferences
  35. Seshadri et al. (2026) Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
  36. Zhu et al. (2024) How Reliable is Your Simulator? Analysis on the Limitations of Current LLM-based User Simulators for Conversational Recommendation
  37. Du et al. (2026) Optimal Design for Active Preference Learning with Biased LLM Judges
  38. Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants

35.3 Preference methods for language-model problems #

The second direction applies tools built for dueling bandits and PBO to problems of language models: optimizing prompts, choosing among outputs and models, choosing which comparisons to collect for training, and exploring online during alignment. The evidence is mixed in a way that turns out to be informative, and it resolves once one asks what each query is for.

35.3.1 Prompt optimization #

APOHF (Lin et al., 2024b) trains a neural network on frozen prompt embeddings with a Bradley-Terry likelihood and pairs the greedy maximizer with the maximizer of an upper confidence bound, the recipe of linear dueling bandits, against baselines that included an ensemble with double Thompson sampling, a dueling-bandit rule that draws two independent posterior samples and pairs their maximizers (Section 21.2). It handles only pairs, and all its "human" feedback was simulated: validation accuracy on 30 instruction tasks and similarity to a target image for text-to-image tasks. It was an ICML 2024 workshop oral, was not accepted at ICLR 2025, and has an ACL Rolling Review submission record from May 2026 with no formal acceptance. Kayal et al. (2025) (ICML 2025) use it as the motivating application of "Bayesian optimization from human feedback". PDO (Wu et al., 2026) (ACL 2026 Findings) schedules duels between prompts with double Thompson sampling under a fixed budget of language-model judgments and uses the winners to guide mutation, finding prompts better than label-free baselines on BIG-bench Hard and MS MARCO. Among the prompt optimizers for text we found, none was evaluated with people; the study of APPO with people concerns images. (Ji, He, and Gu also call an unrelated active-query algorithm APPO.)

35.3.2 Choosing among outputs and among models #

Zhang et al. (2024a) (ICML 2024) replace pointwise scoring of reasoning steps with pairwise comparisons by a language model, handling the noise with ensembles and dueling-bandit variants. Chatbot Arena (Chiang et al., 2024), the public leaderboard built from crowd votes between anonymous models, chooses the next pair in proportion to the expected reduction in the width of the Bradley-Terry confidence intervals; in simulations fitted to 213,576 held-out votes, random sampling needed 54% more data to reach a precision of 0.2 on the win-rate matrix, but only 5% more to reach 0.3 on the scores. Duel-Evolve (Karlekar et al., 2026), an ICLR 2026 workshop poster, fits a Bayesian Bradley-Terry model to a language model's judgments of its own candidates, allocates comparisons by double Thompson sampling, and picks parents for the next generation in evolutionary fashion, with no reward model or labels; it scored 20 percentage points above comparable iterative methods on MathBench and 12 or more on LiveCodeBench, against ground truth. Model routing has been cast as a contextual dueling bandit solved with Feel-Good Thompson sampling, a variant with an optimism term, with lower cumulative regret on RouterBench and MixInstruct (Chiang et al., 2025) (a preprint not accepted at ICLR 2026); Gharat et al. (2026) (NeurIPS 2026) study best-arm identification under dueling feedback with heterogeneous query costs, assuming a Condorcet winner (an option that beats every other with probability above one half) and checking that assumption on real data; CUPID (Nguyen et al., 2026) (ICML 2026) chooses which two models a user should compare, with a study with people; and T-POP (Qu et al., 2026) (ICML 2026) learns one user's reward online with a dueling bandit to steer the decoding of a frozen model.

35.3.3 Active preference data for reward models and DPO #

Choosing which comparisons to collect for training has the most evidence. Muldrew et al. (2024) (ICML 2024) favor pairs on which DPO's implicit preference model is confident but wrong, and with GPT-4 as the oracle improved win rate by 1% to 6% on average. BAL-PM (Melo et al., 2024) (NeurIPS 2024) observes that naive estimates of epistemic uncertainty (uncertainty from lack of data, as opposed to noise in the answers) select redundant samples, adds the entropy of the already-acquired prompts in the model's feature space, and needed 33% to 68% fewer labels on two human preference data sets. Mehta et al. (2025) (COLM 2025) reduce dueling feedback to a contextual Borda function, the probability that an action beats a uniformly random one, prove a polynomial regret bound, and report that their methods can improve performance by over 13% relative to baselines under a limited budget, with uncertainty from Monte Carlo dropout (keeping dropout on at prediction time). Das et al. (2025) (ECML-PKDD 2025) prove that sampling contexts uniformly can leave a constant suboptimality gap under a small budget, give a lower bound of Ω(d/T)\Omega(d/\sqrt{T}) for dd dimensions and TT rounds, and show that APO, which picks the most uncertain context, matches it up to logarithmic and nonlinear factors. Ji et al. (2024) (TMLR) prove a regret of O~(d2/Δ)\tilde O(d^2/\Delta) and a query complexity of O~(d2/Δ2)\tilde O(d^2/\Delta^2), where Δ\Delta is the gap between the best and second-best actions; their ADPO matches DPO with about half as many queries.

A branch based on optimal experimental design uses the Fisher information, the expected curvature of the log-likelihood, whose inverse approximates the covariance of the fitted parameters; a D-optimal design maximizes its determinant, shrinking the uncertainty ellipsoid. It includes an offline design with a lower bound matching up to constant and logarithmic factors (Scheid et al., 2024), D-optimal feedback for a DPO linearized at the last layer (Kveton et al., 2025) (both preprints), a Fisher criterion at the reward model's last layer, with the finding that comparisons across prompts raise labeling efficiency (Shen et al., 2025a) (ICML 2025), a log-determinant choice of negative examples for the Plackett-Luce model (Surana et al., 2026) (a 2026 preprint), and PBO-style acquisition in RLHF with a last-layer Laplace approximation (Section 17.2) (Cercola et al., 2026a) (published in 2026).

ActiveUltraFeedback (Melikidze et al., 2026) (ICML 2026) is the most systematic comparison so far: a contextual dueling bandit over responses from 12 families of open models, an ensemble on a frozen backbone, and a language-model judge's Likert ratings standing in for people. It compared random selection, two heuristics, three classical dueling-bandit rules (infomax, the pair whose choice probability the ensemble disagrees on most; double Thompson sampling; and MaxMinLCB), and two new rules that seek a large quality gap: double reverse Thompson sampling, which pairs the maximizer and minimizer of one posterior sample, and DeltaUCB. Models fine-tuned on 5,000 to 10,000 samples chosen by the new rules outperformed models trained on 60,000 samples chosen by random selection, by UltraFeedback, or by the dueling-bandit rules. For training a reward model, though, random selection was strong, and the authors conclude that there diversity is preferable to a large quality gap.

35.3.4 Thompson sampling for online alignment #

Dwaracherla et al. (2024) (ICML 2024) brought double Thompson sampling and epistemic neural networks, which output a family of predictions indexed by a random input so that its spread expresses what the network does not know, into feedback collection for language models. Queries were a prompt and two Gemini Nano responses; the "human" was a simulator choosing by the Bradley-Terry model from a reward model built on Gemini Pro and fitted to Anthropic's helpfulness and harmlessness data; and the policy was approximated by best-of-NN (keep the highest-scoring of NN responses). Double Thompson sampling beat passive querying, Boltzmann exploration, and infomax, which did well early but then fell far behind, which the authors attribute to its seeking information whether or not it is useful; it reached passive querying's performance with an order of magnitude less data. Asghari et al. (2026), a 2026 preprint, scale this to a 9B Gemma policy, with an epistemic reward network adding fewer than 5% to the parameters, pairs chosen by the variance of the choice probability across ensemble particles (information-directed selection), and a small positive bias on each reinforcement signal, which the authors call an affirmative nudge. With fewer than 20,000 labels it matched what offline RLHF reached with 200,000. The factor of 1,000 in the paper is an extrapolation of curves, the ablations do not separate the selection rule from the nudge, and there is no independent replication.

The theory is well developed: a near-optimal trade-off between regret and queries for trajectory preferences (Wu and Sun, 2024) (ICLR 2024; Section 36.1); FGTS.CDB (Li et al., 2024b) (ICML 2024), the first posterior sampling algorithm for linear contextual dueling bandits, with near minimax optimal regret O~(dT)\tilde O(d\sqrt{T}); neural-tangent-kernel versions of the upper confidence bound and Thompson sampling (Verma et al., 2025) (ICLR 2025); and an O(T)O(\sqrt{T}) bound for Thompson sampling in online RLHF with general function approximation (Feng and Fu, 2025) (a 2025 preprint, without experiments).

Two applied versions followed. SEA (Liu et al., 2024b) frames alignment as a contextual dueling bandit solved with Thompson sampling on models of 1B to 6.9B parameters with DPO, IPO, and SLiC, its preferences coming from a scalar reward model (Skywork-Reward-Llama-3.1-8B) plus one experiment with GPT-4o-mini as judge; it is a NeurIPS 2024 workshop poster, not accepted at ICLR 2025, NeurIPS 2025, or ICLR 2026. warmPref-PS (Agnihotri et al., 2024), a preprint not accepted at TMLR, warm-starts posterior sampling with offline preferences from an expert of unknown competence.

A parallel branch replaces posterior sampling with exploration bonuses: XPO (Xie et al., 2025) (ICLR 2025), value-incentivized preference optimization (Cen et al., 2025) (ICLR 2025), self-exploring language models (Zhang et al., 2024b) (TMLR), and a count-based bonus added to the DPO objective (Bai et al., 2025) (ICLR 2025). Under KL or α\alpha-divergence regularization (a family that contains the KL) these bonuses "unintentionally" bias exploration toward the reference model's high-probability regions (Li et al., 2026d) (ICLR 2026); the sample complexity of all existing online RLHF algorithms grows exponentially with the scale of the reward (Chen et al., 2025a) (NeurIPS 2025); and iterative Nash preference optimization (Section 35.4.4) without explicit exploration can depend exponentially on the KL parameter (Nan et al., 2026) (a 2026 preprint).

35.3.5 Random is hard to beat #

Oh et al. (2026b), an ICLR 2026 workshop paper (I Can't Believe It's Not Better), compared uncertainty-based active preference learning with random selection in online DPO, across harmlessness, helpfulness, and instruction following, with a reward model and a language-model judge as proxies. Active selection "yields negligible improvements in proxy win-rates compared to Random"; while proxy win rate rose, capability on standard benchmarks fell, and active selection neither prevented that nor clearly reduced variance. The authors point to the strong prior from web-scale pretraining and the "cheap diversity" of random on-policy samples. Other results qualify the positive ones: naive uncertainty selects redundant samples (BAL-PM), infomax fell behind after the early phase (Dwaracherla et al.), quality-gap rules beat the classical dueling rules (ActiveUltraFeedback), and Chatbot Arena's savings were 54% for one target and 5% for another. And the positive results simulate people with a reward model (Dwaracherla et al., Asghari et al.), use accuracy or similarity as the utility (APOHF), or use language-model judges (Muldrew et al., ActiveUltraFeedback, PDO, Duel-Evolve); BAL-PM's savings come from human preference data sets, and Chatbot Arena's from simulations fitted to real votes.

35.3.6 What the query is for #

The results can be reconciled once the purpose of a query is separated (inference). The acquisition rules of dueling bandits and PBO (double Thompson sampling, MaxMinLCB, qEUBO, infomax) were designed to find or confirm the best option. Alignment data need pairs that teach a policy or a reward model: DPO training benefits from a large quality gap (ActiveUltraFeedback), reward-model training from diversity (the random-selection results of BAL-PM and ActiveUltraFeedback), and offline methods from coverage of the policy's own distribution (Oh et al., and Song et al. in Section 35.4.3). Using a find-the-best rule unchanged answers a different question from the one being asked. Two further differences matter: alignment picks from a small, policy-dependent candidate set rather than a whole design space, so much of the disagreement between "active" and "random" may come from how diverse that set already is (inference); and Asghari et al. have an explicit posterior over rewards, whereas Oh et al. have only the implicit reward of a directly optimized policy.

As of September 2026 we found no study that compares double Thompson sampling, information-directed selection, quality-gap selection, and random selection under the same budget, with human labels, for both reward-model RLHF and DPO. Until one does, a team collecting preference data should keep a random arm as its baseline and choose the active rule by what the data are for (inference).

Sources cited in Section 35.3 37
  1. Lin et al. (2024b) Prompt Optimization with Human Feedback
  2. Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
  3. Wu et al. (2026) LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
  4. Zhang et al. (2024a) Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
  5. Chiang et al. (2024) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
  6. Karlekar et al. (2026) Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
  7. Chiang et al. (2025) LLM Routing with Dueling Feedback
  8. Gharat et al. (2026) Cost-Aware Best-LLM Identification using Dueling Feedback
  9. Nguyen et al. (2026) CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM
  10. Qu et al. (2026) T-POP: Test-Time Personalization with Online Preference Feedback
  11. Muldrew et al. (2024) Active Preference Learning for Large Language Models
  12. Melo et al. (2024) Deep Bayesian Active Learning for Preference Modeling in Large Language Models
  13. Mehta et al. (2025) Sample Efficient Preference Alignment in LLMs via Active Exploration
  14. Das et al. (2025) Active Preference Optimization for Sample Efficient RLHF
  15. Ji et al. (2024) Reinforcement Learning from Human Feedback with Active Queries
  16. Scheid et al. (2024) Optimal Design for Reward Modeling in RLHF
  17. Kveton et al. (2025) Active Learning for Direct Preference Optimization
  18. Shen et al. (2025a) Active Reward Modeling: Adaptive Preference Labeling for Large Language Model Alignment
  19. Surana et al. (2026) MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization
  20. Cercola et al. (2026a) Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference
  21. Melikidze et al. (2026) ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
  22. Dwaracherla et al. (2024) Efficient Exploration for LLMs
  23. Asghari et al. (2026) Efficient Exploration at Scale
  24. Wu and Sun (2024) Making RL with Preference-based Feedback Efficient via Randomization
  25. Li et al. (2024b) Feel-Good Thompson Sampling for Contextual Dueling Bandits
  26. Verma et al. (2025) Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
  27. Feng and Fu (2025) Thompson Sampling in Online RLHF with General Function Approximation
  28. Liu et al. (2024b) Sample-Efficient Alignment for LLMs
  29. Agnihotri et al. (2024) Online Bandit Learning with Offline Preference Data for Improved RLHF
  30. Xie et al. (2025) Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
  31. Cen et al. (2025) Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
  32. Zhang et al. (2024b) Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
  33. Bai et al. (2025) Online Preference Alignment for Language Models via Count-based Exploration
  34. Li et al. (2026d) General Exploratory Bonus for Optimistic Exploration in RLHF
  35. Chen et al. (2025a) Avoiding $\mathbfexp(R_max)$ scaling in RLHF through Preference-based Exploration
  36. Nan et al. (2026) Efficient Exploration for Iterative Nash Preference Optimization
  37. Oh et al. (2026b) Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs

35.4 A shared likelihood and seven broken analogies #

The two fields share an observation model and, for the online problem, a decision-theoretic framing, so an acquisition rule from one field can run in the other almost unchanged. That shared core is easy to over-read; this section separates it from seven places where the analogy breaks.

35.4.1 What is shared #

The likelihood. Zhu et al. (2023) (ICML 2023) prove that with linear rewards the maximum likelihood estimate converges under both the Bradley-Terry-Luce and Plackett-Luce models but that a policy trained on it can fail, whereas a pessimistic estimate works under a coverage assumption; that for rankings of KK items the full Plackett-Luce estimate is "asymptotically more efficient" than splitting the ranking into pairs; and that RLHF and maximum-entropy inverse reinforcement learning share one analysis. Sun et al. (2025) (ICLR 2025) give convergence rates for Bradley-Terry reward models on deep embeddings and argue that Bradley-Terry is "not a necessary choice", since downstream optimization needs only an order-consistent reward. Tang et al. (2024) (ICML 2024) show that DPO, IPO, and SLiC differ only in their convex loss, which corresponds to the choice of link in PBO. On the PBO side, BoTorch's PairwiseGP (Meta Platforms, Inc., 2026h) (software documentation) and LILO use the probit link of Thurstone (Section 16.3, Section 27.1), while Kayal et al. and PF-TS (Lazzaro et al., 2026) (AISTATS 2026) analyze the Bradley-Terry-Luce model.

The online decision problem. Xiong et al. (2024) (ICML 2024) formalize RLHF as a "reverse-KL regularized contextual bandit" with finite-sample guarantees, and Ji et al., Mehta et al., and SEA adopt the contextual dueling bandit form. Regret guarantees exist on both sides (Kayal et al. and PF-TS for PBO, Section 29.3; Wu and Sun, Ji et al., and APO for RLHF), and acquisition rules transfer almost unchanged.

35.4.2 Where the analogy breaks #

The objective. By Equation (35.5), the optimum of the KL-regularized objective lets the policy network represent both the language model and the implicit reward (Rafailov et al., 2023). Korbak et al. (2022) (EMNLP 2022 Findings) observe that plain reinforcement learning fine-tuning leads to "distribution collapse", whereas KL-regularized reinforcement learning is "equivalent to variational inference", approximating a Bayesian posterior that updates the prior language model with the reward as evidence. PBO returns the maximizer of a latent utility; alignment returns the reference policy tilted exponentially by the reward, and measures regret against that tilted policy. Won et al. (2025), a 2025 preprint, argue that when preferences encode the information needed to update a reference policy into a target, the log-ratio reward is the only reasonable choice.

The implicit reward as a utility estimate. Reading DPO's implicit reward like a Gaussian process posterior mean is not reliable: it fits the training data as well as an explicit reward model but is 3% less accurate on average, and up to 7% less, across five out-of-distribution settings (Lin et al., 2024a) (EMNLP 2024 Findings), and most preference-tuned models rank the pairs of common preference data sets correctly "less than 60%" of the time, with the DPO objective ill-suited to correcting even mild ranking errors of the reference model (Chen et al., 2024) (NeurIPS 2024).

Strong preferences. When preferences are close to deterministic, Bradley-Terry reward differences diverge and the KL regularization grows ever weaker, so DPO overfits, especially when each pair appears only a few times (Azar et al., 2024) (AISTATS 2024). In PBO the Gaussian process prior constrains the latent utility directly, so the problem does not arise in the same form (inference). Table 35.3 lays out all nine aspects.

Table 35.3 PBO and preference-based alignment, aspect by aspect: two analogies that hold and seven that break.
Aspect PBO assumes Preference-based alignment assumes Holds? Why, and the evidence
Observation model probit link (PairwiseGP, LILO) or logistic Bradley-Terry (Kayal et al., PF-TS) logistic Bradley-Terry; Plackett-Luce for lists; DPO, IPO, and SLiC correspond to logistic, squared, and hinge losses holds both are pairwise random utility likelihoods; the links differ only in the noise distribution (Zhu et al. 2023; Tang et al. 2024; Rafailov et al. 2023)
Online decision problem a dueling bandit with cumulative or simple regret a contextual dueling bandit with reverse-KL regularization holds the same acquisition rules apply (Xiong et al. 2024; Ji et al.; Mehta et al. 2025)
1. Objective the maximizer of the latent utility the reference policy tilted exponentially by the reward breaks regret is measured against a tilted reference policy (Rafailov et al. 2023; Korbak et al. 2022)
2. Sample complexity conditional upper bounds: preference feedback of the same order as order-optimal scalar feedback (Kayal et al.) O(1/ε)O(1/\varepsilon) rather than O(1/ε2)O(1/\varepsilon^2) under KL regularization; existing online algorithms exponential in the reward scale; iterative Nash methods can be exponential in the KL parameter breaks the KL term to a reference policy, not the likelihood, drives the rates (Zhao et al. 2025; Chen et al. 2025; Nan et al. 2026)
3. Strong preferences the GP prior constrains the latent utility near-deterministic preferences make reward differences diverge and the KL constraint loses force breaks nothing like a prior on the reward holds the fit in place (Azar et al. 2024)
4. Utility estimate a posterior mean and variance DPO's implicit reward has no uncertainty; 3% less accurate out of distribution on average; ranking accuracy mostly below 60% breaks the implicit reward is a by-product of fitting a policy (Lin et al. 2024; Chen et al. 2024)
5. Query distribution any pair anywhere in the design space candidates sampled from the policy or a model pool; offline methods need global coverage; exploration bonuses drift toward the reference model's high-probability region breaks what can be compared is limited to what the policy generates (Song et al. 2024; Li et al. 2026)
6. Scale tens to hundreds of comparisons, about 2 to 20 dimensions, Laplace or expectation propagation 10410^4 to 10610^6 comparisons over sequences; ensembles, dropout, or Laplace breaks a summary across the studies in this chapter (inference)
7. Respondents one decision maker with a consistent utility many annotators or language-model judges breaks a single reward fitted to many people raises problems of social choice (Gölz et al. 2025; Siththaranjan et al. 2024; Chidambaram et al. 2026)

Three of these differences, sample size, heterogeneity of the respondents, and where the queries come from, can be seen by fitting the same likelihood in both worlds, and Figure 35.2 also shows the first, what each world returns.

One person · 50 comparisons · one utility · pairs from anywhereOne likelihood in both worlds: P(a ≻ b) = sigmoid(r(a) − r(b))fitted reward95% bandthe person's utility−202utility or rewardwhere comparisons fallreturned design0.00.20.40.60.81.0option: a design, or a response on one axisReward sd where compared: 0.60Error against the person's utility, where compared: 0.43Returns one design: x = 0.80 (the person's best: 0.75)
One person · 50 comparisonsone utility · pairs from anywhereOne likelihood in both worlds: P(a ≻ b) = sigmoid(r(a) − r(b))fitted reward95% bandthe person's utility−202rewardwhere comparisons fallreturned design0.00.20.40.60.81.0optionReward sd where compared: 0.60Error against the person's utility: 0.43Returns one design: x = 0.80 (the person's best: 0.75)
Figure 35.2 One Bradley-Terry likelihood, P(a ≻ b) = sigmoid(r(a) − r(b)), fitted to simulated comparisons in two worlds, with options on one axis. The fit uses a Gaussian prior on 15 radial basis functions (a small stand-in for a Gaussian process) and a Laplace band. One person: 50 comparisons from one utility, pairs drawn uniformly anywhere, standing in for an acquisition function free to query anywhere. Many annotators: a million comparisons from two groups with different favorites, pairs drawn from a reference policy that never produces options in the shaded regions. The strip shows where comparisons fall and what each world returns: one design (the best fitted option among those that can be shown) or the reference policy tilted by exp(r/β), the optimum of KL-regularized alignment (Equation (35.4)), here with β = 1. Utilities, policy, and β are illustrative; PBO's most-used implementation uses the closely related probit link (Section 27.5).

Things to try:

  1. Start in the one-person world and press New draw a few times. With 50 comparisons the band is wide and the returned design moves around the person's favorite near 0.75. Drag NN to 1,000,000: the band collapses onto the person's utility.
  2. Raise the second group's share to 0.4. The band stays thin, but the fitted reward matches neither group, and the error no longer shrinks as NN grows: one reward fitted to two groups is a compromise (Section 35.4.4).
  3. Switch Where pairs come from to From the policy. Outside the region the policy covers, the band stays wide however large NN grows: no amount of data reaches options the policy never generates.
  4. Switch the world to Many annotators. The orange curve, the reference policy reweighted by er/βe^{r/\beta}, can only reweight options the policy produces and so never reaches group A's favorite near 0.75. Back in One person with pairs from the policy and N=50N = 50, PBO restricted to someone else's candidates stops at the edge of what the policy produces (Section 35.3.6).

35.4.3 Sample complexity and coverage #

Write ε\varepsilon for how close to optimal the learned policy must be. A sample complexity of O(1/ε2)O(1/\varepsilon^2) means that halving ε\varepsilon needs four times the data; O(1/ε)O(1/\varepsilon) means twice. Zhao et al. (2025) (NeurIPS 2025) point out that earlier analyses of KL-regularized RLHF gave the same O(1/ε2)O(1/\varepsilon^2) as the unregularized problem, and that a sharper analysis gives O(1/ε)O(1/\varepsilon). Song et al. (2024) (NeurIPS 2024) prove that a global coverage condition, roughly that the data contain every response the optimal policy might produce, is necessary and sufficient for DPO-like offline contrastive methods to reach the optimal policy, while online reinforcement learning needs only partial coverage; step 3 above shows what a lack of coverage looks like. The main difference between the theories of PBO and RLHF therefore lies in the KL constraint to a reference policy and in the policy parameterization, not in the preference likelihood (inference).

35.4.4 Many annotators: from optimization to social choice #

The difference in respondents pushes alignment toward social choice, the study of how to combine many people's preferences into one decision (Section 40.7), which PBO has barely started on (Section 20.5). With cyclic preferences there may be no Condorcet winner. The Borda count scores each option by how often it beats an opponent drawn at random; a von Neumann winner is a probability distribution over options that beats or ties every single option in expectation, and exists even when preferences cycle; a Nash equilibrium of a two-player game is a pair of strategies neither player can improve on alone. Nash learning from human feedback (Munos et al., 2024) (ICML 2024) seeks a policy preferred to any opponent, the Nash equilibrium of a two-player constant-sum game, and with arbitrary preferences the target of preference-based reinforcement learning becomes the von Neumann winner (Wang et al., 2023a) (NeurIPS 2023). The von Neumann winner is also a solution concept of the dueling-bandit literature, so dueling-bandit results on intransitive preferences (Section 21.1) are the natural counterpart of Nash learning on the PBO side (inference); we did not check whether a paper from 2025 or 2026 makes this connection formally.

Gölz et al. (2025) (NeurIPS 2025) measure alignment methods by their distortion, the worst-case ratio between the best achievable average utility and that of the learned policy, and prove that Nash learning achieves minimax-optimal distortion while RLHF and DPO can have exponential or unbounded distortion in the full setting; Oko et al. (2026) (ICML 2026) then prove that under reward clipping the exponential degradation comes from a mismatch between the preference data and the reference policy, not from the algorithm. The dispute is open. What a single fitted reward does with many people has also been characterized: with unobserved context, such as which annotator answered, a single learned utility implicitly aggregates by Borda count (Siththaranjan et al., 2024) (ICLR 2024), as does the Bradley-Terry-Luce loss (An et al., 2026) (a 2026 preprint), which is the compromise curve of step 2 above; with limited data per user, binary comparisons cannot identify latent user types whereas rankings of three or more items can (Chidambaram et al., 2026) (AISTATS 2026); and under Bayesian marginalization or KL-robust optimization the effective reward of the KL-regularized objective has a closed form, optimistic in the Bayesian branch and pessimistic in the robust one (Hahami et al., 2026) (a 2026 preprint). The same reward posterior should then be used optimistically when choosing queries and pessimistically when choosing what to deploy, and the acquisition functions of PBO handle only the first (inference). Section 41.3 takes up what alignment to many people can mean.

Sources cited in Section 35.4 22
  1. Zhu et al. (2023) Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons
  2. Sun et al. (2025) Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
  3. Tang et al. (2024) Generalized Preference Optimization: A Unified Approach to Offline Alignment
  4. Meta Platforms, Inc. (2026h) BoTorch PairwiseGP source code pairwise_gp.py
  5. Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
  6. Xiong et al. (2024) Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
  7. Rafailov et al. (2023) Direct Preference Optimization: Your Language Model is Secretly a Reward Model
  8. Korbak et al. (2022) RL with KL penalties is better viewed as Bayesian inference
  9. Won et al. (2025) Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
  10. Lin et al. (2024a) On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
  11. Chen et al. (2024) Preference Learning Algorithms Do Not Learn Preference Rankings
  12. Azar et al. (2024) A General Theoretical Paradigm to Understand Learning from Human Preferences
  13. Zhao et al. (2025) Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
  14. Song et al. (2024) The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
  15. Munos et al. (2024) Nash Learning from Human Feedback
  16. Wang et al. (2023a) Is RLHF More Difficult than Standard RL?
  17. Gölz et al. (2025) Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
  18. Oko et al. (2026) Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner
  19. Siththaranjan et al. (2024) Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
  20. An et al. (2026) Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences
  21. Chidambaram et al. (2026) Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
  22. Hahami et al. (2026) A Unifying Lens on Reward Uncertainty in RLHF

35.5 Common claims, checked #

Table 35.4 checks common claims about how PBO relates to alignment.

Table 35.4 Common claims about PBO and language-model alignment, checked against the evidence.
Claim What the evidence says
PBO and RLHF share the Bradley-Terry model. Partly. Both use pairwise random utility likelihoods, but the default PBO implementation, PairwiseGP, uses the probit link. Logistic forms appear in the original PBO paper of González et al. (González et al., 2017), in POP-BO (Xu et al., 2024b) (ICML 2024), in MaxMinLCB, and in MR-LPF (the algorithm of Kayal et al.).
DPO shares a probit or Bradley-Terry likelihood with PBO. "Probit" is wrong. The DPO loss is the logistic Bradley-Terry log-likelihood, with a Plackett-Luce version in the appendix.
Bradley-Terry is the Rosetta stone connecting BO, RLHF, and DPO; DPO's implicit reward is conceptually like a Gaussian process posterior. True at the level of the likelihood, and only an explanatory analogy for the implicit reward, which carries no uncertainty (Table 35.3).
RLHF scales but is query-inefficient; PBO is query-efficient but does not scale; a hybrid gets the best of both. The direction holds, with qualifications. Hybrids exist (an ensemble posterior with double Thompson sampling or information-directed selection), but their posteriors are neural ensembles, not Gaussian processes, and the same kind of method gave negligible gains over random selection in online DPO. We found no work that uses a Gaussian process or kernel surrogate as the reward model in an RLHF loop at language-model scale.
Bringing PBO's uncertainty into RLHF, and warm-starting PBO with priors from language models, are open challenges. Both have been pursued since 2024: the first by Dwaracherla et al., BAL-PM, and Cercola et al., with positive and negative results; the second by OPEN, MAPLE, and Eichelbeck et al., all but OPEN with simulated users.
Ji et al.'s RLHF with active queries appeared at NeurIPS. The content is right (ADPO reaches DPO's level with about half the queries), but the venue is TMLR.
RLHF reward models are "notoriously poorly calibrated". We could not substantiate this as stated, and it usually comes without a citation. Thies et al. (2026a) (ICML 2026) found that the calibration of RLHF reward models "correlates strongly but not perfectly with benchmark accuracy".
PBO uses about 10 to 1,000 comparisons and RLHF about 50,000 to 500,000 (or: tens to one or two hundred, against thousands to millions). The contrast in order of magnitude holds, but neither range has a source. Studies of PBO with people use tens to one or two hundred comparisons; simulated high-dimensional PBO uses about 10 to 15 per dimension, about 1,500 at 102 dimensions (Menn et al., 2026a) (a 2026 preprint); alignment uses 10410^4 to 10610^6.
LiPO shows that the gains from listwise preference data grow monotonically with list length. Partly. Only the LiPO-λ loss benefits monotonically from longer lists; a Plackett-Luce form of DPO does not (Liu et al., 2025a) (NAACL 2025).
PEBOL raises MAP@10 by up to 131% after 10 turns. That figure (mean average precision at 10) appears only in the first arXiv version (Austin et al., 2024b). The RecSys version reports mean reciprocal rank at 10 instead: at most 0.27, against 0.17 for the best monolithic baseline.
Sources cited in Section 35.5 6
  1. González et al. (2017) Preferential Bayesian Optimization
  2. Xu et al. (2024b) Principled Preferential Bayesian Optimization
  3. Thies et al. (2026a) Calibrated Preference Learning: The Case of Label Ranking
  4. Menn et al. (2026a) Local Preferential Bayesian Optimization
  5. Liu et al. (2025a) LiPO: Listwise Preference Optimization through Learning-to-Rank
  6. Austin et al. (2024b) Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation

35.6 What flows back #

The main direction of transfer has been from bandits and PBO into alignment: double Thompson sampling, epistemic neural networks, information-directed sampling, kernelized dueling bandits, and uncertainty-aware reward models. Transfer back has so far been mostly framing, applications, and evaluation. MaxMinLCB (Pásztor et al., 2024) (NeurIPS 2024), a kernelized preference bandit that casts the choice of a pair as a zero-sum Stackelberg game (one player commits first and the other responds), motivates itself by noting that such a model "has been employed in systems for fine-tuning large language models", and its group then published RewardUQ (Yang et al., 2026b), a EurIPS 2025 workshop paper, and ActiveUltraFeedback. Kayal et al. use prompt optimization as their application; PF-TS draws its two competitors symmetrically, which its authors argue is often desirable, and sometimes necessary, with human evaluators or language-model judges; and the multi-user dueling bandit of Ahmed and Ghasemi (2026) (TMLR) opens with the unfairness to minority groups of training on average preferences in language-model fine-tuning, and proves a lower bound on regret for fairness across DD users of Ω(T2/3min⁡(K,D)1/3)\Omega(T^{2/3}\min(K, D)^{1/3}) with KK arms. Ax 1.2.3 added an abstraction for language-model messages, and 1.3.0 language-in-the-loop labeling trials and qEUBO scheduling for PBO (Meta Platforms, Inc., 2026l) (software documentation). Not everything that looks like a transfer is one: LILO extends BOPE (Lin et al., 2022) (AISTATS 2022), which predates the alignment wave, and the abstracts of PABBO (Zhang et al., 2025a) (ICLR 2025) and POP-BO do not frame their problems in terms of language models.

35.6.1 Three results ready to transfer #

Three results from alignment could be used in PBO directly. We found no PBO paper that takes DPO, LiPO, PAL, or an uncertainty-aware reward model as the source of a method component, though we did not systematically scan the reference lists of PBO papers from 2024 to 2026 (inference).

Rankings carry more than pairs. Zhu et al. proved that the full Plackett-Luce estimate beats splitting rankings into pairs, and Chidambaram et al. that rankings of three items identify latent user types that pairs cannot. Katkuri et al. (2026), a 2026 preprint, learn rewards with Plackett-Luce from a vision-language model's rankings and match or beat pairwise Bradley-Terry and RL-VLM-F (Wang et al., 2024b) (ICML 2024) on Meta-World; LiPO (Liu et al., 2025a) applies learning-to-rank losses (Section 36.6) to listwise preference optimization; and GraphDPO (Liu et al., 2026a), a 2026 preprint, generalizes preferences to a graph. The PBO interfaces that ask for choices among several options, GimmBO (Liu et al., 2026b) (SIGGRAPH North America 2026) and MultiBO (Rajagopalan et al., 2026) (ICML 2026), adopted lists for reasons of interface design (Section 20.1) and cite none of these formal arguments.

Separating noise from ignorance. Several alignment papers split reward uncertainty into an aleatoric part, the irreducible noise in the answers, and an epistemic part, uncertainty from lack of data; Lou et al. (2024), a 2024 preprint, use a probabilistic value head for the first and ensemble disagreement for the second. Others quantify reward uncertainty to mitigate over-optimization, with reward-model ensembles (Coste et al., 2024) (ICLR 2024) or a Laplace approximation on LoRA weights, a small set of low-rank adapter weights (Yang et al., 2024) (a workshop paper), and BNRM adds non-negative factor analysis to the Bradley-Terry model (Duan et al., 2026) (ICML 2026); RewardUQ found that model size and initialization matter more than the choice of uncertainty method. Human comparison noise is aleatoric, while acquisition by disagreement needs the epistemic part.

Modeling the oracle's bias. NAOD puts the judge's bias into acquisition (Section 35.2.4); LILO, which uses a language-model oracle, does not.

35.6.2 Population priors from pluralistic reward models #

Pluralistic and personalized reward models could serve as population priors for a new user of PBO. PAL (Chen et al., 2025c) (ICLR 2025) uses an ideal-point model, in which each user and option is a point in a shared latent space and preference falls with distance, with mixture modeling, and generalizes to new users from a few samples. The variational preference learning of Poddar et al. (2024) (NeurIPS 2024) infers a latent vector per user and suffers posterior collapse (the per-user latent stops carrying information) when each user has little data (Kim and Kim, 2026) (ICLR 2026). LoRe (Bose et al., 2025) (COLM 2025) and Cai et al. (2026) (SIGIR 2026) represent each user's reward as a weighted combination of shared basis functions, and on a benchmark of personalized reward models the best reached 75.94% accuracy (Ma et al., 2026) (COLM 2026). None chooses queries to learn a user's weights, so active per-user elicitation on a meta-learned low-rank prior is the natural next step for PBO to take from them (inference).

ActiveUltraFeedback's result also points to a distinction PBO has not yet drawn explicitly: the pairs that help locate the optimum are not the pairs that help learn a reusable utility model (inference). PBO work that builds population priors from earlier users is in the second situation; HOMI, for example, pretrains a prior from user models and was significantly better only at the second and third iterations (Liao et al., 2026) (CHI 2026).

Sources cited in Section 35.6 23
  1. Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
  2. Yang et al. (2026b) RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
  3. Ahmed and Ghasemi (2026) Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare
  4. Meta Platforms, Inc. (2026l) CHANGELOG (versions 1.2 to 1.3)
  5. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  6. Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
  7. Katkuri et al. (2026) Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning
  8. Wang et al. (2024b) RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
  9. Liu et al. (2025a) LiPO: Listwise Preference Optimization through Learning-to-Rank
  10. Liu et al. (2026a) Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph
  11. Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
  12. Rajagopalan et al. (2026) Personalized Image Generation via Human-in-the-loop Bayesian Optimization
  13. Lou et al. (2024) Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
  14. Coste et al. (2024) Reward Model Ensembles Help Mitigate Overoptimization
  15. Yang et al. (2024) Bayesian Reward Models for LLM Alignment
  16. Duan et al. (2026) Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
  17. Chen et al. (2025c) PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences
  18. Poddar et al. (2024) Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
  19. Kim and Kim (2026) Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
  20. Bose et al. (2025) LoRe: Personalizing LLMs via Low-Rank Reward Modeling
  21. Cai et al. (2026) One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
  22. Ma et al. (2026) Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
  23. Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models

35.7 Settled, contested, missing #

Research status Settled, contested, missing

Settled. PBO and preference-based alignment share pairwise random utility likelihoods (Bradley-Terry, Thurstone, Plackett-Luce) and the contextual dueling-bandit framing, and acquisition rules move between them almost unchanged. The optimum of the KL-regularized objective is the reference policy tilted by er/βe^{r/\beta}, and DPO follows because the normalizer cancels in Bradley-Terry (Rafailov et al., 2023). The default PBO implementation uses the probit link, DPO the logistic one. Language models used as simulated users and judges are close to people in aggregate but unreliable for individuals, and biased by position, length, and self-preference. Language models acting directly as dueling-bandit optimizers trail classical algorithms in strong regret (Xia et al., 2025), and in scalar BO, agents performed no differently when their observations were replaced by random labels (Gupta et al., 2025). Only two of the eight systems in Table 35.1 are PBO in the strict sense.

Contested. Whether active selection of comparisons beats random selection for alignment: gains of 1% to 6% in win rate, 33% to 68% fewer labels, and an order of magnitude in simulated settings, against negligible gains in online DPO (Oh et al., 2026b); the reconciliation by the purpose of the query is our inference. Whether the exponential distortion of RLHF and DPO belongs to the algorithms (Gölz et al., 2025) or to a mismatch between data and reference policy (Oko et al., 2026). Whether the large label savings of Asghari et al. hold, given ablations that do not separate their two components and no replication. Whether the Bradley-Terry form is needed at all (Sun et al., 2025).

Missing. A PBO loop run with human comparisons and with language-model comparisons, reporting regret or sample efficiency for both. A measurement of how a judge's position bias propagates into acquisition decisions. A comparison, with human labels and a common budget, of double Thompson sampling, information-directed selection, quality-gap selection, and random selection for both reward-model RLHF and DPO. An evaluation with people of any preference-based prompt optimizer for text. A Gaussian process or kernel surrogate as the reward model of an RLHF loop at language-model scale. PBO methods built from alignment components (Plackett-Luce rankings, the aleatoric and epistemic split, oracle-bias models, pluralistic priors with active per-user elicitation), and an acquisition function that distinguishes pairs for finding the optimum from pairs for learning a reusable utility.

Sources cited in Section 35.7 7
  1. Rafailov et al. (2023) Direct Preference Optimization: Your Language Model is Secretly a Reward Model
  2. Xia et al. (2025) Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
  3. Gupta et al. (2025) LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
  4. Oh et al. (2026b) Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
  5. Gölz et al. (2025) Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
  6. Oko et al. (2026) Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner
  7. Sun et al. (2025) Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Further reading #

References

  1. Agnihotri, A., Jain, R., Ramachandran, D., and Wen, Z. (2024). Online Bandit Learning with Offline Preference Data for Improved RLHF. arXiv (not accepted at TMLR). preprint Cited in §35.3
  2. Ahmed, M. H., and Ghasemi, M. (2026). Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare. Transactions on Machine Learning Research. Cited in §35.6
  3. An, Z., Nakshbandi, D., and Du, W. (2026). Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences. arXiv. preprint Cited in §35.4
  4. Asghari, S. M., Chute, C., Dwaracherla, V., Lu, X., Jafarnia, M., Minden, V., Wen, Z., and Van Roy, B. (2026). Efficient Exploration at Scale. arXiv. preprint Cited in §35.3
  5. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Cited in §35.2
  6. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024b). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. arXiv. preprint Cited in §35.5
  7. Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R. (2024). A General Theoretical Paradigm to Understand Learning from Human Preferences. AISTATS 2024. Cited in §35.1 §35.4
  8. Bai, C., Zhang, Y., Qiu, S., Zhang, Q., Xu, K., and Li, X. (2025). Online Preference Alignment for Language Models via Count-based Exploration. ICLR 2025. Cited in §35.3
  9. Bose, A., Xiong, Z., Chi, Y., Du, S. S., Xiao, L., and Fazel, M. (2025). LoRe: Personalizing LLMs via Low-Rank Reward Modeling. Conference on Language Modeling (COLM 2025). Cited in §35.6
  10. Cai, H., Li, Y., Yu, T., Zhu, F., Wang, W., Feng, F., and Li, W. (2026). One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment. SIGIR 2026. Cited in §35.6
  11. Cen, S., Mei, J., Goshvadi, K., Dai, H., Yang, T., Yang, S., … Dai, B. (2025). Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF. ICLR 2025. Cited in §35.3
  12. Cercola, M., Capretti, V., and Formentin, S. (2026a). Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference. IFAC Journal of Systems and Control. doi:10.1016/j.ifacsc.2026.100398. Cited in §35.3
  13. Chawla, A., Thompson, W. H. W., and Young, J.-G. (2026). Multiple latent orderings better predict language model preferences. arXiv. preprint Cited in §35.2
  14. Chen, A., Malladi, S., Zhang, L. H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K. (2024). Preference Learning Algorithms Do Not Learn Preference Rankings. Advances in Neural Information Processing Systems. doi:10.52202/079017-3234. Cited in §35.4
  15. Chen, M., Chen, Y., Sun, W., and Zhang, X. (2025a). Avoiding scaling in RLHF through Preference-based Exploration. Advances in Neural Information Processing Systems 38 (NeurIPS 2025). doi:10.52202/085713-5485. Cited in §35.3
  16. Chen, D., Chen, Y., Rege, A., and Vinayak, R. K. (2025c). PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences. ICLR 2025. Cited in §35.6
  17. Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., … Stoica, I. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. ICML 2024. Cited in §35.3
  18. Chiang, C.-K., Ishida, T., and Sugiyama, M. (2025). LLM Routing with Dueling Feedback. arXiv (not accepted at ICLR 2026). preprint Cited in §35.3
  19. Chidambaram, K., Seetharaman, K. V., and Syrgkanis, V. (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Cited in §35.4
  20. Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems. Cited in §35.1
  21. Coste, T., Anwar, U., Kirk, R., and Krueger, D. (2024). Reward Model Ensembles Help Mitigate Overoptimization. International Conference on Learning Representations. Cited in §35.1 §35.6
  22. Das, N., Chakraborty, S., Pacchiano, A., and Chowdhury, S. R. (2025). Active Preference Optimization for Sample Efficient RLHF. Machine Learning and Knowledge Discovery in Databases. Research Track. doi:10.1007/978-3-032-06096-9_6. Cited in §35.3
  23. Du, Z., Zhang, H., Zhu, H., and Zhang, B. (2026). Optimal Design for Active Preference Learning with Biased LLM Judges. arXiv. preprint Cited in §35.2
  24. Duan, Z., Rong, G., Li, Z., Chen, B., Zhou, M., and Guo, D. (2026). Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling. ICML 2026. Cited in §35.6
  25. Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., … Hashimoto, T. B. (2023). AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback. Advances in Neural Information Processing Systems. doi:10.52202/075280-1308. Cited in §35.2
  26. Dwaracherla, V., Asghari, S. M., Hao, B., and Van Roy, B. (2024). Efficient Exploration for LLMs. ICML 2024. Cited in §35.3
  27. Eichelbeck, M., Voigt, T., and Althoff, M. (2026). Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent Space. International Conference on Learning Representations. Cited in §35.2
  28. Feng, S., and Fu, J. (2025). Thompson Sampling in Online RLHF with General Function Approximation. arXiv. preprint Cited in §35.3
  29. Gharat, S., Karamchandani, N., and Nair, J. (2026). Cost-Aware Best-LLM Identification using Dueling Feedback. Advances in Neural Information Processing Systems. Cited in §35.3
  30. Gölz, P., Haghtalab, N., and Yang, K. (2025). Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? NeurIPS 2025. Cited in §35.4 §35.7
  31. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §35.5
  32. Gupta, R., Hartford, J., and Liu, B. (2025). LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Findings of the Association for Computational Linguistics: EMNLP 2025. Cited in §35.2 §35.7
  33. Hahami, E., Zimmermann, Y., Zhou, R., and Benarroch Jedlicki, J. (2026). A Unifying Lens on Reward Uncertainty in RLHF. arXiv. preprint Cited in §35.4
  34. Handa, K., Gal, Y., Pavlick, E., Goodman, N., Andreas, J., Tamkin, A., and Li, B. Z. (2024). Bayesian Preference Elicitation with Language Models. arXiv. preprint Cited in §35.2
  35. Ji, K., He, J., and Gu, Q. (2024). Reinforcement Learning from Human Feedback with Active Queries. TMLR. Cited in §35.3
  36. Karlekar, S., Zheng, C., Saebo, M., Beltran-Velez, N., Yu, S., Bowlan, J., Kucer, M., and Blei, D. (2026). Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. ICLR 2026 RSI Workshop. workshop paper Cited in §35.3
  37. Katkuri, S., Kawada, M., and Wachs, J. (2026). Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning. arXiv. preprint Cited in §35.6
  38. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §35.3
  39. Kim, G., and Kim, E. (2026). Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback. ICLR 2026. Cited in §35.6
  40. Kirk, H. R., Leqi, L., Zeng, F., Davidson, H., Vidgen, B., Summerfield, C., and Hale, S. A. (2026). PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users. arXiv. preprint Cited in §35.2
  41. Kobalczyk, K., Astorga, N., Liu, T., and van der Schaar, M. (2025). Active Task Disambiguation with LLMs. ICLR 2025. Cited in §35.2
  42. Kobalczyk, K., Lin, Z. J., Letham, B., Zhao, Z., Balandat, M., and Bakshy, E. (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Cited in §35.2
  43. Korbak, T., Perez, E., and Buckley, C. L. (2022). RL with KL penalties is better viewed as Bayesian inference. Findings of the Association for Computational Linguistics: EMNLP 2022. doi:10.18653/v1/2022.findings-emnlp.77. Cited in §35.4
  44. Kristiadi, A., Strieth-Kalthoff, F., Skreta, M., Poupart, P., Aspuru-Guzik, A., and Pleiss, G. (2024a). A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules? International Conference on Machine Learning. Cited in §35.2
  45. Kuric, E., Demcak, P., and Krajcovic, M. (2026). Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design? arXiv. preprint Cited in §35.2
  46. Kveton, B., Li, X., McAuley, J., Rossi, R., Shang, J., Wu, J., and Yu, T. (2025). Active Learning for Direct Preference Optimization. arXiv. preprint Cited in §35.3
  47. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §35.4
  48. Li, X., Zhao, H., and Gu, Q. (2024b). Feel-Good Thompson Sampling for Contextual Dueling Bandits. International Conference on Machine Learning. Cited in §35.3
  49. Li, B. Z., Tamkin, A., Goodman, N., and Andreas, J. (2025b). Eliciting Human Preferences with Language Models. ICLR 2025. Cited in §35.2
  50. Li, W., Oh, C., and Li, S. (2026d). General Exploratory Bonus for Optimistic Exploration in RLHF. International Conference on Learning Representations. Cited in §35.3
  51. Li, Z., Liao, Y.-C., and Holz, C. (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. Cited in §35.2
  52. Liao, Y.-C., Belo, J., Moon, H.-S., Steimle, J., and Feit, A. M. (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Cited in §35.6
  53. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §35.2 §35.6
  54. Lin, Y., Seto, S., ter Hoeve, M., Metcalf, K., Theobald, B.-J., Wang, X., … Zhang, T. (2024a). On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization. Findings of the Association for Computational Linguistics: EMNLP 2024. doi:10.18653/v1/2024.findings-emnlp.940. Cited in §35.4
  55. Lin, X., Dai, Z., Verma, A., Ng, S.-K., Jaillet, P., and Low, B. K. H. (2024b). Prompt Optimization with Human Feedback. ICML 2024 MHFAIA Workshop (no formal proceedings). workshop paper Cited in §35.3
  56. Liu, T., Astorga, N., Seedat, N., and van der Schaar, M. (2024a). Large Language Models to Enhance Bayesian Optimization. International Conference on Learning Representations. Cited in §35.2
  57. Liu, Z., Chen, C., Du, C., Lee, W. S., and Lin, M. (2024b). Sample-Efficient Alignment for LLMs. NeurIPS 2024 LanGame Workshop. workshop paper Cited in §35.3
  58. Liu, T., Qin, Z., Wu, J., Shen, J., Khalman, M., Joshi, R., … Wang, X. (2025a). LiPO: Listwise Preference Optimization through Learning-to-Rank. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). doi:10.18653/v1/2025.naacl-long.121. Cited in §35.5 §35.6
  59. Liu, N., Sun, C., Klinkner, K., and Malmasi, S. (2026a). Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph. arXiv. preprint Cited in §35.6
  60. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §35.6
  61. Lou, X., Yan, D., Shen, W., Yan, Y., Xie, J., and Zhang, J. (2024). Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown. arXiv (withdrawn from ICLR 2025). preprint Cited in §35.6
  62. Ma, Q., Gao, D., Cai, R., Zhao, B., Zhou, H., Zhang, J., and Zhao, Z. (2026). Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization. COLM 2026. Cited in §35.6
  63. Mahmud, S., Nakamura, M., and Zilberstein, S. (2025). MAPLE: A Framework for Active Preference Learning Guided by Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v39i26.34964. Cited in §35.2
  64. Mehta, V., Belakaria, S., Das, V., Neopane, O., Dai, Y., Bogunovic, I., … Neiswanger, W. (2025). Sample Efficient Preference Alignment in LLMs via Active Exploration. COLM 2025. Cited in §35.3
  65. Melikidze, D., Schneider, M., Lam, J., Wertich, M., Hakimi, I., Pásztor, B., and Krause, A. (2026). ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning. ICML 2026. Cited in §35.3
  66. Melo, L. C., Tigas, P., Abate, A., and Gal, Y. (2024). Deep Bayesian Active Learning for Preference Modeling in Large Language Models. NeurIPS 2024. Cited in §35.3
  67. Menn, J., Kober, M., Brunzema, P., Stenger, D., and Trimpe, S. (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Cited in §35.5
  68. Meta Platforms, Inc. (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Cited in §35.4
  69. Meta Platforms, Inc. (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Cited in §35.6
  70. Muldrew, W., Hayes, P., Zhang, M., and Barber, D. (2024). Active Preference Learning for Large Language Models. ICML 2024. Cited in §35.2 §35.3
  71. Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., … Piot, B. (2024). Nash Learning from Human Feedback. ICML 2024. Cited in §35.4
  72. Nan, T., Li, X., Kroer, C., and Lin, T. (2026). Efficient Exploration for Iterative Nash Preference Optimization. arXiv. preprint Cited in §35.3
  73. Nguyen, S., Liu, X., and Senanayake, R. (2026). CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM. International Conference on Machine Learning (ICML 2026). Cited in §35.3
  74. Niwa, R., Yoshida, S., Koyama, Y., and Ushiku, Y. (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Cited in §35.2
  75. Oh, G., Lee, J., Park, J., Yu, Y., Bae, W., and Noh, J. (2026b). Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs. ICLR 2026 Workshop: I Can't Believe It's Not Better (ICBINB). workshop paper Cited in §35.3 §35.7
  76. Oko, K., Ulichney, A., Haghtalab, N., and Bao, H. (2026). Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner. International Conference on Machine Learning. Cited in §35.4 §35.7
  77. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems. Cited in §35.1
  78. Panickssery, A., Bowman, S. R., and Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS 2024. Cited in §35.2
  79. Pásztor, B., Kassraie, P., and Krause, A. (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Cited in §35.6
  80. Peng, Y.-H., Bigham, J. P., and Wu, J. (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Cited in §35.2
  81. Piriyakulkij, W. T., Kuleshov, V., and Ellis, K. (2023). Active Preference Inference using Language Models and Probabilistic Reasoning. NeurIPS 2023 FMDM Workshop. workshop paper Cited in §35.2
  82. Poddar, S., Wan, Y., Ivison, H., Gupta, A., and Jaques, N. (2024). Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning. Advances in Neural Information Processing Systems 37 (NeurIPS 2024). doi:10.52202/079017-1664. Cited in §35.6
  83. Qiu, L., Sha, F., Allen, K., Kim, Y., Linzen, T., and van Steenkiste, S. (2026). Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models. Nature Communications. doi:10.1038/s41467-025-67998-6. Cited in §35.2
  84. Qu, Z., Zhang, M., Kong, M., Li, X., Shang, Z., Wang, Z., … Dai, Z. (2026). T-POP: Test-Time Personalization with Online Preference Feedback. International Conference on Machine Learning. Cited in §35.3
  85. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. Cited in §35.1 §35.4 §35.7
  86. Rajagopalan, R., Dutta, D., Wei, Y.-L., and Roy Choudhury, R. (2026). Personalized Image Generation via Human-in-the-loop Bayesian Optimization. International Conference on Machine Learning. Cited in §35.6
  87. Ramos, M. C., Michtavy, S. S., Porosoff, M. D., and White, A. D. (2026). Bayesian Optimization of Catalysis With In-Context Learning. ACS Central Science. doi:10.1021/acscentsci.5c02418. Cited in §35.2
  88. Ranković, B., Griffiths, R.-R., and Schwaller, P. (2026). Large language models as uncertainty-calibrated optimizers for experimental discovery. Nature Machine Intelligence. doi:10.1038/s42256-026-01283-z. Cited in §35.2
  89. Rodrigues, C., Vas, O., DCosta, I. A., and Prabhakaran, N. K. (2026). When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model. arXiv. preprint Cited in §35.2
  90. Rychert, A., Spagnolo, G., and Posashkov, E. (2025). Reproducibility Study of Large Language Model Bayesian Optimization. arXiv. preprint Cited in §35.2
  91. Saracay, I., Schmidt, L., and Guestrin, C. (2026). Beyond expert users: agents should help users construct preferences, not just elicit them. Conference on Language Modeling (COLM 2026). Cited in §35.2
  92. Scheid, A., Boursier, E., Durmus, A., Jordan, M. I., Ménard, P., Moulines, E., and Valko, M. (2024). Optimal Design for Reward Modeling in RLHF. arXiv. preprint Cited in §35.3
  93. Schoinas, E., Rastogi, A., Carter, A., Granley, J., and Beyeler, M. (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Cited in §35.2
  94. Seshadri, P., Cahyawijaya, S., Odumakinde, A., Singh, S., and Goldfarb-Tarrant, S. (2026). Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2026.acl-long.2192. Cited in §35.2
  95. Shen, Y., Sun, H., and Ton, J.-F. (2025a). Active Reward Modeling: Adaptive Preference Labeling for Large Language Model Alignment. International Conference on Machine Learning. Cited in §35.3
  96. Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., and Vosoughi, S. (2025). Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics. Cited in §35.2
  97. Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Cited in §35.4
  98. Song, Y., Swamy, G., Singh, A., Bagnell, J. A., and Sun, W. (2024). The Importance of Online Data: Understanding Preference Fine-tuning via Coverage. NeurIPS 2024. Cited in §35.4
  99. Sun, H., Shen, Y., and Ton, J.-F. (2025). Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives. ICLR 2025. Cited in §35.4 §35.7
  100. Surana, R., Li, X., Yu, S., Shen, Y. J., Wang, C., Yu, T., … Wu, J. (2026). MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization. arXiv. preprint Cited in §35.3
  101. Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., … Piot, B. (2024). Generalized Preference Optimization: A Unified Approach to Offline Alignment. International Conference on Machine Learning. Cited in §35.1 §35.4
  102. Thies, S. M. A. R., Bengs, V., Kaufmann, T., Vollmer, S. J., and Hüllermeier, E. (2026a). Calibrated Preference Learning: The Case of Label Ranking. International Conference on Machine Learning (ICML 2026). Cited in §35.5
  103. Verma, A., Dai, Z., Lin, X., Jaillet, P., and Low, B. K. H. (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. Cited in §35.3
  104. Wang, Y., Liu, Q., and Jin, C. (2023a). Is RLHF More Difficult than Standard RL? NeurIPS 2023. Cited in §35.4
  105. Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., … Sui, Z. (2024a). Large Language Models are not Fair Evaluators. ACL 2024. Cited in §35.2
  106. Wang, Y., Sun, Z., Zhang, J., Xian, Z., Biyik, E., Held, D., and Erickson, Z. (2024b). RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback. International Conference on Machine Learning. Cited in §35.6
  107. Won, Y., Lee, H., Hwang, H., and Seo, M. (2025). Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization. arXiv. preprint Cited in §35.4
  108. Wu, R., and Sun, W. (2024). Making RL with Preference-based Feedback Efficient via Randomization. ICLR 2024. Cited in §35.3
  109. Wu, Y., Verma, S., Lee, J., Xiong, F., Zhang, P., Awadelkarim, A., … Hill, S. (2026). LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization. Findings of the Association for Computational Linguistics: ACL 2026. doi:10.18653/v1/2026.findings-acl.490. Cited in §35.3
  110. Xia, F., Liu, H., Yue, Y., and Li, T. (2025). Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of the Association for Computational Linguistics: ACL 2025. doi:10.18653/v1/2025.findings-acl.519. Cited in §35.2 §35.7
  111. Xie, T., Foster, D. J., Krishnamurthy, A., Rosset, C., Awadallah, A., and Rakhlin, A. (2025). Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF. ICLR 2025. Cited in §35.3
  112. Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T. (2024). Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint. ICML 2024. Cited in §35.4
  113. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §35.5
  114. Xu, Y., Ruis, L., Rocktäschel, T., and Kirk, R. (2025a). Investigating Non-Transitivity in LLM-as-a-Judge. International Conference on Machine Learning. Cited in §35.2
  115. Yang, A. X., Robeyns, M., Coste, T., Shi, Z., Wang, J., Bou-Ammar, H., and Aitchison, L. (2024). Bayesian Reward Models for LLM Alignment. ICLR 2024 SeT LLM Workshop; ICML 2024 SPIGM Workshop. workshop paper Cited in §35.6
  116. Yang, J., Hu, Z., Qiu, C., Deng, Z., Jiao, X., and Zhou, T. (2026a). Quantifying and Mitigating Self-Preference Bias of LLM Judges. arXiv. preprint Cited in §35.2
  117. Yang, D., Stante, S., Redhardt, F., Libon, L., Kassraie, P., Hakimi, I., Pásztor, B., and Krause, A. (2026b). RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models. EurIPS 2025 EIML Workshop. workshop paper Cited in §35.6
  118. Yuan, X., Chen, Z., Zhang, J., Xiong, H., Ye, N., Li, Y., and Gu, Q. (2026). Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery. ICLR 2026. Cited in §35.2
  119. Zhang, Z.-Y., Han, S., Yao, H., Niu, G., and Sugiyama, M. (2024a). Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought. International Conference on Machine Learning. Cited in §35.3
  120. Zhang, S., Yu, D., Sharma, H., Zhong, H., Liu, Z., Yang, Z., … Wang, Z. (2024b). Self-Exploring Language Models: Active Preference Elicitation for Online Alignment. TMLR. Cited in §35.3
  121. Zhang, X., Huang, D., Kaski, S., and Martinelli, J. (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §35.6
  122. Zhao, H., Ye, C., Gu, Q., and Zhang, T. (2025). Sharp Analysis for KL-Regularized Contextual Bandits and RLHF. NeurIPS 2025. Cited in §35.4
  123. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., … Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems. doi:10.52202/075280-2020. Cited in §35.2
  124. Zhu, B., Jordan, M., and Jiao, J. (2023). Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. International Conference on Machine Learning. Cited in §35.4
  125. Zhu, L., Huang, X., and Sang, J. (2024). How Reliable is Your Simulator? Analysis on the Limitations of Current LLM-based User Simulators for Conversational Recommendation. Companion Proceedings of the ACM Web Conference 2024. doi:10.1145/3589335.3651955. workshop paper Cited in §35.2