Bayesian Optimization
Part X: Synthesis
中文

What a Comparison Measures

Preferential Bayesian optimization (PBO) looks for the design a person likes best by asking them to compare options. A statistical model reads each answer as a noisy comparison of an unobserved number attached to every design, the person's utility: by default a Gaussian process prior (a distribution over smooth functions) on the utility function, and a probit link that turns the difference between two utilities into a choice probability (Chapter 16, Chapter 18). An acquisition function picks the next pair (Chapter 19), and after a few dozen to a couple of hundred answers the design with the highest estimated utility is the result.

That model assumes one fixed utility function and noise of constant size. This chapter gathers what the earlier parts found about that assumption to answer a plain question: what does one answer to "which do you prefer?" measure? Our answer is four things at once, so a PBO session is both an estimate of a preference and an intervention on it. A sentence that is our own inference ends with "(inference)".

45.1 The bottleneck moved #

Between 2017 and 2026, most of the field's effort went into acquisition functions and regret theory, and both advanced. The bottleneck has since moved from algorithms to measurement: what a single comparison measures, how answers should be modeled, and what asking does to the person who answers. The largest errors now sit in the observation model, the part of the model that says how a person's answer arises from their utility, and in how methods are evaluated against real people.

45.1.1 Acquisition: from heuristics to decision theory #

The acquisition rules common from 2017 to 2021 adapted scalar Bayesian optimization (Section 28.1); four independent groups reported that one of them, modified expected improvement, stalls (Section 28.4), and on a constructed instance batch expected improvement is not asymptotically consistent: even unlimited queries need not find the best option (Astudillo et al., 2023). The rule that replaced these heuristics is EUBO, the expected utility of the best option, which scores a pair by the expected utility of whichever member the person would pick (Section 19.4); qEUBO is its multi-option form. EUBO is one-step Bayes optimal, the best possible choice if only one question remained, a property with an older root in decision analysis: the optimal set of options to recommend is also the myopically optimal set to ask about (Viappiani and Boutilier, 2020). Its flaws were documented in 2026, in two preprints: its queries collapse toward the estimated maximum, and its pairs tend to form isolated pieces of the comparison graph, which leaves the Hessian of the Laplace approximation rank deficient (Wu and Gardner, 2026; Shao et al., 2026). Section 19.6 tells both stories, and Section 19.5.1 shows a similar collapse in a recorded simulation.

Rankings of acquisition functions flip with the metric, the noise model, and the inference method: in the paper that introduced the optimistic algorithm POP-BO, qEUBO reported a slightly better final solution while its cumulative regret, the summed shortfall of the options shown along the way (Chapter 13), was more than 2.5 times higher (Xu et al., 2024b). qEUBO with BoTorch's PairwiseGP model is the best-maintained default implementation (Section 27.5), a standing that comes from the software ecosystem, not from comparisons with simpler methods.

45.1.2 Theory: mature, conditional, and only upper bounds #

Regret theory for preference feedback matured, but every result carries conditions, and all are upper bounds: guarantees that regret grows no faster than some rate in the number of queries TT. Two quantities recur. γT\gamma_T, the maximum information gain of the kernel, measures how much TT noisy observations can reveal about a function drawn from the Gaussian process prior (Section 6.5). κ\kappa is an upper bound on the inverse slope of the link function, which becomes large when choice probabilities saturate near 0 or 1 (Section 29.5). The sequential algorithms MaxMinLCB and PF-TS have preference-probability regret O~(γTT)\tilde O(\gamma_T \sqrt{T}), where O~\tilde O ignores logarithmic factors, with constants that grow with κ\kappa (Pásztor et al., 2024; Lazzaro et al., 2026), and a batched algorithm does better only under extra conditions (Kayal et al., 2025) (Section 29.4). There is no kernelized lower bound under a logistic or probit link (Section 29.7). And the theory covers elimination, optimistic, and Thompson algorithms on frequentist kernel estimators, while practice runs a Laplace posterior with EUBO, for which there are only one-step optimality and finite-domain consistency results; nothing connects the two (inference).

Whether a comparison costs more than a number depends on the goal. For the regret of a single fixed utility, the upper bounds in the finite-arm, linear, and kernel settings show that pairwise comparisons are not more expensive in order. For identifying a heterogeneous population, pairwise data are the weakest feedback: a model fitted to binary comparisons implicitly aggregates by Borda count, which scores each option by its average chance of beating the others (Siththaranjan et al., 2024), and rankings of at least three options are needed to identify latent user types (Chidambaram et al., 2026).

45.1.3 The observation model is the main bottleneck #

The probit link with constant noise remains the default, and its extensions mostly come from one research group with one evaluation each (Chapter 27). Human answers depart a long way from "a fixed utility plus noise of constant size". When 20 people with design training judged the same 600 pairs of generated interfaces, their agreement was 0.25 on Krippendorff's α\alpha, a chance-corrected agreement score on which 0 is chance and 1 is perfect agreement, in a 2026 preprint (Peng et al., 2026). Among 35 chemists, Fleiss' kappa, a score of the same kind, was 0.40 and 0.32 in two rounds (Choung et al., 2023). The same moral pairwise questions, repeated within and across sessions, had on average 6% to 20% of answers flip (Keswani et al., 2026). In retinal-implant optimization, sighted participants and the simulated agents meant to stand in for them agreed on only about 50% of choices (Schoinas et al., 2025). And in a meta-analysis of the stability of risk preference, the estimated reliability (the correlation between two measurements of the same people taken close together) was 0.61 for self-reported propensity to take risks and only 0.25 for behavioral measures, standardized tasks such as choices between lotteries (Bagaïni et al., 2025). A pairwise choice is a behavioral measure, so a few stated-preference questions alongside the comparisons might carry the stable component (inference). Figure 45.1 puts these numbers side by side.

chance, or no stable signalperfect consistencyDifferent people, same pairsDesigners, generated interfaces0 (chance)1 (perfect)0.25Chemists, molecules0 (chance)1 (perfect)0.32 and 0.40Same person, same question againMoral questions, asked again50% (coin flip)0% (never flips)6% to 20%Simulated stand-in against the personRetinal implant, simulated agents50% (chance)100% (always)about 50%Same people measured twice (risk preference)Self-reported risk taking0 (no stable signal)1 (perfect)0.61Behavioral choices (lotteries)0 (no stable signal)1 (perfect)0.25Designers, generated interfaces: Krippendorff's α, 0.25Scale from 0 (chance) to 1 (perfect); the value sits 25% of the way along it.20 people with design training judged the same 600 pairs of generated interfaces (Peng andcolleagues, a 2026 preprint).
chance, or no stable signalperfect consistencyDifferent people, same pairsDesigners, generated interfaces0 (chance)1 (perfect)0.25Chemists, molecules0 (chance)1 (perfect)0.32 and 0.40Same person, same question againMoral questions, asked again50% (coin flip)0% (never flips)6% to 20%Simulated stand-in against the personRetinal implant, simulated agents50% (chance)100% (always)about 50%Same people measured twice (risk preference)Self-reported risk taking0 (no stable signal)1 (perfect)0.61Behavioral choices (lotteries)0 (no stable signal)1 (perfect)0.25Designers, generated interfaces:Krippendorff's α, 0.25Scale from 0 (chance) to 1 (perfect); the valuesits 25% of the way along it.20 people with design training judged thesame 600 pairs of generated interfaces (Pengand colleagues, a 2026 preprint).
Figure 45.1 How consistent human answers are, in the studies cited in this section. Each row is a different statistic on its own scale, stretched so that chance (or no stable signal) sits at the left end and perfect consistency at the right; the rows can be read for where they fall between those ends, not compared with each other in size. The flip rate is drawn reversed, since fewer flips mean more consistency. Choose a row to read its study and statistic.

Some things to try:

  • Choose the flips row. A flip rate of 6% to 20% means the same person repeats 80% to 94% of their answers; a person answering by coin flip would repeat 50%.
  • Compare the two reliability rows. They come from one meta-analysis and use one statistic, so they can be compared: self-reports (0.61) are more than twice as reliable as behavioral measures (0.25).

Simulation studies of acquisition functions assume an observation model, so they cannot measure the error that comes from that model being wrong; improving the likelihood is more likely to change research conclusions than another acquisition function (inference). Of the signals that cost the person no extra effort, response times (Shvartsman et al., 2024; Li et al., 2024a) and stated confidence (Zhang et al., 2026b) have improved models on real human data, so far in offline fits rather than in closed-loop optimization.

45.1.4 Small samples, weak evaluation #

Studies of pure preference feedback in interactive design mostly enroll 6 to 60 people, applied studies elsewhere 1 to 35 evaluators with a median below 10, and validation is almost always internal: the same person later chooses the result (Chapter 32, Section 34.4). Benchmarks are weak in the same direction: scalar test functions in 1 to 8 dimensions, at least six noise models and five definitions of regret, no shared benchmark, and no public data set of individual-level pairwise judgments (Chapter 31). Even an expert's own cost function can miss their choices: in a robot-commissioning study with a single expert operator, reported in a 2025 preprint, a cost function the expert had designed did not fully capture the expert's own choices, even after its weights were refitted to the expert's highest-rated experiments (De Witte et al., 2025). Synthetic decision makers defined by a hidden cost function may therefore overstate how well methods work on real people (inference).

45.1.5 Simpler methods often do as well #

In several settings, simpler methods perform comparably: people tuning an exoskeleton themselves reached, in about 11 minutes, metabolic savings of the same order as algorithmic tuning (Schäfer et al., 2026); a spherical input mapping with Bayesian linear regression reached the state of the art on high-dimensional benchmarks (Doumont et al., 2026); and random selection was hard to beat in online direct preference optimization of language models, in a 2026 workshop paper (Oh et al., 2026b), while the authors of the amortized optimizer PABBO write that a random strategy often beats some Gaussian process baselines (Zhang et al., 2025a). Theory points to where PBO's advantage must come from: when ranking items actively from noisy comparisons, parametric assumptions such as Bradley-Terry or Thurstone buy at most a logarithmic gain (Heckel et al., 2019), so PBO's sample efficiency should come mainly from the kernel sharing information between neighboring designs, not from the link function (inference).

The setting in which PBO remains supported by evidence is therefore narrow: options that can be judged only by perceptual comparison, no numerical objective, a dimension already low after the representation has been designed, and a budget of tens to about two hundred comparisons. Section 46.9 turns this into advice, including the baselines a study needs before it can claim an advantage for PBO.

Key idea Where the error is

Acquisition functions now have a decision-theoretic footing and regret bounds have caught up with scalar feedback, but the error that matters most comes from what the model assumes about the person, which simulations with a known observation model cannot see (inference).

Sources cited in Section 45.1 24
  1. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  2. Viappiani and Boutilier (2020) On the equivalence of optimal recommendation sets and myopically optimal query sets
  3. Wu and Gardner (2026) Knowledge Gradient for Preference Learning
  4. Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
  5. Xu et al. (2024b) Principled Preferential Bayesian Optimization
  6. Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
  7. Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
  8. Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
  9. Siththaranjan et al. (2024) Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
  10. Chidambaram et al. (2026) Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
  11. Peng et al. (2026) Efficient Personalization of Generative User Interfaces
  12. Choung et al. (2023) Extracting medicinal chemistry intuition via preference machine learning
  13. Keswani et al. (2026) Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
  14. Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
  15. Bagaïni et al. (2025) A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures
  16. Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences
  17. Li et al. (2024a) Enhancing Preference-based Linear Bandits via Human Response Time
  18. Zhang et al. (2026b) Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
  19. De Witte et al. (2025) How to Capture Human Preference: Commissioning of a Robotic Use-Case via Preferential Bayesian Optimisation
  20. Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
  21. Doumont et al. (2026) We Still Don't Understand High-Dimensional Bayesian Optimization
  22. Oh et al. (2026b) Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
  23. Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
  24. Heckel et al. (2019) Active ranking from pairwise comparisons and when parametric assumptions do not help

45.2 What the disciplines changed #

Part IX collected results about preference from psychology, economics, neuroscience, and other fields. Since 2017, some failed in preregistered replications (studies whose hypotheses and analyses are fixed before data collection) or shrank in meta-analyses corrected for publication bias; others were strengthened and gained mechanisms that can be written into a likelihood. Table 45.1 and Table 45.2 list only the results with direct consequences for preferential optimization; the last column of each is our inference. The effect sizes dd and gg express a difference between conditions in standard deviations, so d=0.06d = 0.06 is a twentieth of a standard deviation and d=0.40d = 0.40 is a moderate effect.

Table 45.1 Results about preference that weakened after 2017, and what that means for PBO (last column: inference).
Result Status Key evidence What it means for PBO (inference)
Ego depletion and domain-general decision fatigue: deciding draws on a limited resource weakened to near zero 36 laboratories, d=0.06d = 0.06 (Vohs et al., 2021); 231,076 triage calls show no decision fatigue (Andersson et al., 2025) Do not set session length on this basis; measure fatigue within the task
Induced-compliance dissonance: people change attitudes to match what they were induced to do not replicated 39 laboratories, N=4,898N = 4{,}898 (Vaidis et al., 2024) Explain post-choice preference change by revaluation, not by classic dissonance
Moral licensing within a person: a good deed licenses a later lapse near zero after bias correction; only a reputational effect when observed remains gg between −0.08-0.08 and −0.02-0.02 (Rotella et al., 2026) Do not model compensation between successive answers in private sessions
Mood as information: current mood is read as evidence about the options much weakened most of nine replications not significant (Yap et al., 2017) Do not schedule exploration by mood
Value discrimination worsens near the optimum fails for value judgments; holds for perceptual parameters decisions between high-value options are faster and more accurate (Shevlin et al., 2022) Value comparisons may become more accurate near convergence; perceptual parameters keep a just-noticeable-difference floor
A loss-aversion coefficient of about 2.25 value unstable meta-analytic mean 1.955 (Brown et al., 2024); about 1.07 when symmetric and unsorted (Yechiam and Zeif, 2025) Do not fix an asymmetric likelihood around the current best
The attraction (decoy) effect: adding a dominated option boosts its neighbor limited to attributes shown as numbers Frederick et al. (2014) The mechanism is weak in perceptual design
Table 45.2 Results about preference that strengthened or appeared after 2017, and what that means for PBO (last column: inference).
Result Status Key evidence What it means for PBO (inference)
Choice-induced preference change: choosing an option raises its value for the chooser strengthened d=0.40d = 0.40 (Enisman et al., 2021); each choice adds about $0.18 to the chosen item (Zylberberg et al., 2024) Add a choice-history term to the likelihood; the most-compared current best is affected most
Range normalization: values are coded relative to the range on offer strengthened Bavard et al. (2018); Bavard and Palminteri (2023) The utility's scale is session-relative; recalibrate before reusing it across sessions
Efficient coding of value noise: precision goes to the values a person expects to see new and strengthened Polanía et al. (2019); Prat-Carrabin and Woodford (2022) Noise scales with the values presented, so fixed noise misreads late-session precision
Random utility consistency: repeated choices behave as if drawn from a distribution over utilities strengthened for most people; part of the observed intransitivity is genuine most of 141 participants satisfy random utility (McCausland et al., 2020); with response times separating noise from preference, 19.24% and 39.58% of the transitivity violations in two data sets are genuine, and the rest could be either (Alós-Ferrer et al., 2023); the noise specification changes the inference (Bhatia and Loomes, 2017) A scalar utility plus noise works for most people, but the link and noise form are substantive modeling choices
Incomplete preferences and deliberate randomization strengthened, with limits 40% to 50% of participants report incompleteness when allowed, on average 3.3 of 50 comparisons (Nielsen and Rigotti, 2026); about half choose inconsistently with complete preferences plus certainty independence, an axiom about mixing options with a sure outcome (Cettolin and Riedl, 2019); most choose to randomize (Agranov and Ortoleva, 2025) Offer a "can't compare" answer and model it separately from ties and noise
Information per comparison at most one bit, but the error rate is of the same order as with numerical ratings Shah et al. (2016) The structure of the comparison graph, not the bits per query, governs the error

Two of these studies are working papers (Alós-Ferrer et al., 2023; Nielsen and Rigotti, 2026). The rows are developed in Part IX: fatigue in Section 37.2.1, moral licensing in Section 38.5, mood in Section 38.4, loss aversion in Section 40.1, context effects in Section 37.2.3, choice-induced change in Section 37.2.2, range normalization in Section 37.3.1, efficient coding in Section 39.4, random utility in Section 37.4.1, and incomplete preferences in Section 40.4 and Section 43.4.2.

The changes follow a pattern (inference). What weakened were mostly arguments for session-design rules (set session length by decision fatigue, schedule queries by mood) and for fixed numbers for directional biases (a loss-aversion coefficient, compensation in moral licensing). What strengthened were mechanisms with a magnitude that map onto parts of a likelihood. The evidence therefore supports a few mechanisms whose sign and size are estimated in each task, not one likelihood term for every documented bias.

Formal tools PBO has not yet used. Five families of results can be applied to PBO, and as of September 2026 we found no PBO paper that uses them:

  1. comparison-graph theory, in which the comparison graph's Laplacian spectrum and effective resistances govern the error of pairwise estimates (Shah et al., 2016; Hendrickx et al., 2019), and the Hodge decomposition splits comparison data into a part a scalar utility can explain and a cyclic part none can (Jiang et al., 2011; Strang et al., 2022) (Section 43.3.1, Section 43.3.2);
  2. performative prediction, which separates learning a target from steering it (Perdomo et al., 2020; Hardt and Mendler-Dünner, 2025) (Section 42.2);
  3. trust regions on induced preference shift (Carroll et al., 2022), needed because when preferences move toward what a system shows, low regret can be achieved trivially by moving them (Dean and Morgenstern, 2022);
  4. behavioral welfare economics, which calls one option better for a person's welfare only if the other is never chosen over it in any frame (Bernheim and Rangel, 2009) and measures welfare losses from misunderstood consequences (Ambuehl et al., 2022);
  5. multi-utility representations of incomplete preferences, in which an option counts as better only if every utility in a set agrees (Evren and Ok, 2011), with identification results that distinguish indifference, indecisiveness, and experimentation (Ok and Tserenjigmid, 2022).

The first and last apply to PBO logs and posteriors without new theory, so they are the place to start (inference). Relatedly, fitting one Bradley-Terry utility to several people's pooled comparisons is itself an aggregation rule, and it violates Pareto optimality and pairwise majority consistency (Ge et al., 2024).

Sources cited in Section 45.2 34
  1. Vohs et al. (2021) A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect
  2. Andersson et al. (2025) No evidence for decision fatigue using large-scale field data from healthcare
  3. Vaidis et al. (2024) A Multilab Replication of the Induced-Compliance Paradigm of Cognitive Dissonance
  4. Rotella et al. (2026) Observation Moderates the Moral Licensing Effect: A Meta-Analytic Test of Interpersonal and Intrapsychic Mechanisms
  5. Yap et al. (2017) The effect of mood on judgments of subjective well-being: Nine tests of the judgment model
  6. Shevlin et al. (2022) High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity
  7. Brown et al. (2024) Meta-analysis of Empirical Estimates of Loss Aversion
  8. Yechiam and Zeif (2025) Loss aversion is not robust: A re-meta-analysis
  9. Frederick et al. (2014) The Limits of Attraction
  10. Enisman et al. (2021) Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm
  11. Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
  12. Bavard et al. (2018) Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences
  13. Bavard and Palminteri (2023) The functional form of value normalization in human reinforcement learning
  14. Polanía et al. (2019) Efficient coding of subjective value
  15. Prat-Carrabin and Woodford (2022) Efficient coding of numbers explains decision bias and noise
  16. McCausland et al. (2020) Testing the Random Utility Hypothesis Directly
  17. Alós-Ferrer et al. (2023) Identifying Nontransitive Preferences
  18. Bhatia and Loomes (2017) Noisy preferences in risky choice: A cautionary note
  19. Nielsen and Rigotti (2026) Revealed Incomplete Preferences
  20. Cettolin and Riedl (2019) Revealed preferences under uncertainty: Incomplete preferences and preferences for randomization
  21. Agranov and Ortoleva (2025) Ranges of Randomization
  22. Shah et al. (2016) Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
  23. Hendrickx et al. (2019) Graph Resistance and Learning from Pairwise Comparisons
  24. Jiang et al. (2011) Statistical ranking and combinatorial Hodge theory
  25. Strang et al. (2022) The Network HHD: Quantifying Cyclic Competition in Trait-Performance Models of Tournaments
  26. Perdomo et al. (2020) Performative Prediction
  27. Hardt and Mendler-Dünner (2025) Performative Prediction: Past and Future
  28. Carroll et al. (2022) Estimating and Penalizing Induced Preference Shifts in Recommender Systems
  29. Dean and Morgenstern (2022) Preference Dynamics Under Personalized Recommendations
  30. Bernheim and Rangel (2009) Beyond Revealed Preference: Choice-Theoretic Foundations for Behavioral Welfare Economics *
  31. Ambuehl et al. (2022) Evaluating Deliberative Competence: A Simple Method with an Application to Financial Choice
  32. Evren and Ok (2011) On the multi-utility representation of preference relations
  33. Ok and Tserenjigmid (2022) Indifference, indecisiveness, experimentation, and stochastic choice
  34. Ge et al. (2024) Axioms for AI Alignment from Human Feedback

45.3 Found or made? #

Behind these results lies an old question: do repeated comparisons find a preference that was there before, or do they partly make it? The constructive view holds that preferences are built during elicitation, so the measurement changes what it measures; the stable view holds that there is an underlying preference that careful measurement can recover. Preferential optimization takes a side whether its users notice or not. The evidence since 2017 leaves part of each view standing and supports neither extreme.

45.3.1 What each view keeps #

The constructive view keeps its central claim: elicitation changes the object it measures, by a considerable amount and through mechanisms that can be modeled. Choosing an option raises its value (Table 45.2); for 59 people rating 55 morphed dog images, individual models with a learning component explained on average 17% more variance for the real presentation order than for a simulated random order (Brielmann et al., 2024); and forced choice manufactures some inconsistency that would otherwise show up as hesitation (Costa-Gomes et al., 2022). It loses some supporting arguments: the advantage of unconscious thought was not found in a meta-analysis and large replication (Nieuwenstein et al., 2015), the induced-compliance and mood effects failed or shrank (Table 45.1), and Ruth Chang's work on hard choices, sometimes read as calling a forced choice between options "on a par" a category error, treats such a choice as an occasion for commitment (Chang, 2024).

The stable view keeps the claim that there is a component worth estimating. Most people's repeated choices satisfy random utility (McCausland et al., 2020); behavioral biases are nearly unchanged at the population level over three years, in a working paper (Stango and Zinman, 2024); and some anomalies in risky choice look like errors in valuing complex options rather than preferences (Oprea, 2024), though that conclusion is disputed. It loses its strongest inference, that iterated querying converges to the true underlying preference. Plott's "discovered preference" hypothesis has been observed only after forced trading or after people reflected on axioms they themselves endorse (Nielsen and Rehbeck, 2022; Engelmann and Hollard, 2010), not as a natural outcome of repeated comparison, and in the field PBO loops often end before they converge, most evaluation sequences of a three-month deployment stopping at the first iteration (Ou et al., 2022) (Section 46.6).

45.3.2 The hierarchical view #

One position in this debate, which we call the hierarchical view, holds that preferences over ultimate goals (comfort, safety, beauty) always exist, while preferences over intermediate outcomes, such as a particular parameter setting, may have to be elicited. As a description of where stability lives, the evidence largely supports it: stability concentrates at the broad, abstract, self-reported level, and in a longitudinal study in early adolescence, 75% of respondents had value hierarchies that correlated at least 0.85 across two years (Vecchione et al., 2020).

As a prediction about PBO, the view fares worse. The value computed during a choice is relative to the current goal: goal congruence explains choices and their neural correlates better than reward value (Frömer et al., 2019), so a stable ultimate goal need not become a stable utility over pairwise comparisons (inference). Iterated querying has not been seen to converge (Section 45.3.1), and because stability is measured with self-reports and construction with choices, Section 47.1 describes the study that would test both levels on one design task. Nor does rational inattention, the theory that people spend costly attention optimally, single out the logit link as the view's defenders sometimes claim: it yields the logit only under a Shannon-entropy cost of information, and it can generate any additive random utility model (Fosgerau et al., 2020).

45.3.3 Preference scaffolding #

A second position, preference scaffolding, holds that PBO is not a readout of a fixed utility function but a structure that helps a person form a preference, so that evaluation should cover the exploration process and the person's reflective endorsement of the result, not only the final design. The core holds up, with three qualifications. Not all inconsistency is preference formation: in repeated discrete choice tasks, instability concentrates on options close in utility and on hard tasks, and the underlying preferences of unstable respondents do not differ from those of stable ones (Fraser et al., 2021). Scaffolding is not neutral: when preferences can be influenced, each of the eight notions of alignment compared in one analysis either errs toward undesirable influence or is overly risk-averse (Carroll et al., 2024), so helping and steering have to be told apart by conditions the objective does not supply (Section 45.5). And endorsement after the fact cannot by itself justify a change the system caused, because the changed person may endorse the result only because their values were changed (Pettigrew, 2023).

One argument offered for scaffolding does not hold. Active inference, a theory in which action minimizes surprise relative to preferred observations, is sometimes said to dissolve the dispute between construction and revelation, but it only re-encodes it: in Bayesian optimization it reduces to weighted information-gain acquisition functions whose weights still need tuning (Millidge et al., 2021; Li et al., 2026c) (Section 39.5). The most direct support is a 2025 measurement: when Bayesian optimization led the search, designers reached better results but reported significantly less sense of agency, and collaboration through explicit constraints matched collaboration through natural language in performance while giving more agency (Niwa et al., 2025). That supports a scaffold led by the designer, not preference shaping led by the system.

Sources cited in Section 45.3 19
  1. Brielmann et al. (2024) Modelling individual aesthetic judgements over time
  2. Costa-Gomes et al. (2022) Choice, deferral, and consistency
  3. Nieuwenstein et al. (2015) On making the right choice: A meta-analysis and large-scale replication attempt of the unconscious thought advantage
  4. Chang (2024) What’s so Hard about Hard Choices?
  5. McCausland et al. (2020) Testing the Random Utility Hypothesis Directly
  6. Stango and Zinman (2024) Behavioral Biases Are Temporally Stable
  7. Oprea (2024) Decisions under Risk Are Decisions under Complexity
  8. Nielsen and Rehbeck (2022) When Choices Are Mistakes
  9. Engelmann and Hollard (2010) Reconsidering the Effect of Market Experience on the "Endowment Effect"
  10. Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
  11. Vecchione et al. (2020) Stability and change of basic personal values in early adolescence: A 2‐year longitudinal study
  12. Frömer et al. (2019) Goal congruency dominates reward value in accounting for behavioral and neural correlates of value-based decision-making
  13. Fosgerau et al. (2020) Discrete Choice and Rational Inattention: A General Equivalence Result
  14. Fraser et al. (2021) Preference stability in discrete choice experiments. Some evidence using eye-tracking
  15. Carroll et al. (2024) AI Alignment with Changing and Influenceable Reward Functions
  16. Pettigrew (2023) Nudging for changing selves
  17. Millidge et al. (2021) Whence the Expected Free Energy?
  18. Li et al. (2026c) Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference
  19. Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction

45.4 Four components of a comparison #

Putting the pieces together, we propose that a single pairwise answer mixes four components, whose proportions are moderated by two further factors (Table 45.3; inference). No single paper has tested this decomposition, and the proportions change with familiarity with the domain, the similarity of the options, and the length of the session, so they have to be estimated in each application rather than assumed.

Table 45.3 Four components of a pairwise answer and two moderating factors. The decomposition is our inference; the evidence column says where each piece of support comes from.
Component What it is Evidence, and where it comes from Representation in the model
Stable preference an estimable latent utility tests of random utility, stability of biases, value hierarchies; from lotteries, money choices, and self-reports, not from design comparisons the latent utility of a Gaussian process or other surrogate
Structured evaluation noise noise that varies with difficulty, similarity, protocol, ambiguity of the representation, and the distribution of values presented early versus late noise (Shen et al., 2025b); efficient coding; range normalization; mechanisms supported by experiments, only offline fits in PBO heteroscedastic noise; a joint likelihood with response times; separate noise scales per feedback type (Ghosal et al., 2023)
Query-induced change revaluation after choosing, anchoring on the system's proposals, familiarity and fatigue choice-induced preference change; biases grow after interacting with a biased AI (Glickman and Sharot, 2025); magnitudes from free-choice paradigms, never measured in design comparisons a choice-history term; a drift term; balanced query order as a control
Incomplete or deliberately randomized answers the person cannot or will not compare, or deliberately randomizes experiments on incomplete and randomized preferences; from lotteries and money choices; the share in design comparisons is unknown a "can't compare" answer; a mixture parameter; a multi-utility representation
Moderator 1: how much taste is shared the informativeness of a population prior varies by domain shared taste is high for faces and landscapes and low for architecture and artworks; where landscapes and exterior architecture were compared within the same people, the judgments were equally reliable (Vessel et al., 2018); individual models against the group average, median correlation 0.65 against 0.01 (Brielmann et al., 2024) the weight of a population prior, as a domain or personal parameter
Moderator 2: legitimacy conditions normative conditions that separate helping to form a preference from steering it many candidate conditions, no consensus, none tested in PBO not part of the likelihood; part of the experimental design and the report

The first component is the one the default model assumes. The second says that noise is not one number: it grows with the ambiguity of how the options are represented (early noise) and with time pressure (late noise) (Shen et al., 2025b), depends on the values a person has recently seen (Polanía et al., 2019; Bavard and Palminteri, 2023), and near the optimum can make candidates perceptually indistinguishable, as in a 2026 preprint on prosthesis tuning (Taddei et al., 2026) (Table 46.8). The third component has no place in the default model at all. The fourth is invisible under forced choice, where a person with no preference must still answer and the answer looks like noise.

Which component dominates depends on the domain (inference). Familiar domains and broad tendencies come closer to noisy discovery of a stable preference; novel, similar, multi-attribute options and long sessions come closer to construction and drift. PBO's candidates are usually close variants of one design, and in repeated aesthetic ratings the assimilation and contrast effects between successive items grow with the similarity of the stimuli (Pombo et al., 2023), so the mix has to be measured in each application.

45.4.1 A simulated session #

Within a single session the four components are hard to tell apart, because each can produce mostly consistent answers and a posterior that sharpens. Figure 45.2 shows this in a simulation whose components, forms, and numbers are illustrative choices of ours. A simulated person compares designs on a line. Their stable preference has two peaks of different heights. Structured noise has a standard deviation that rises from 0.15 for very different designs toward 0.15 plus the slider value for nearly identical ones. Query-induced change raises the chosen design by the slider value after every choice and lowers the rejected one by half that; most of it fades over the next ten answers (15% of what is left with each answer), but the lasting share remains. Incomplete answers are a coin flip under forced choice, or "can't compare" when it is offered. The model is the default of Chapter 19, with pairs chosen by EUBO or at random.

stable preferencethe person nowmodel's fitted utilityAcquisition order (EUBO) · one week later−0.6−0.4−0.20.00.20.40.6utilityearly favoritefinal designrejected earlyAnswers, first at top● chosen · ○ rejected · grey: no preference behind itfirst 100.00.20.40.60.81.0designRetest without the system: does the person still prefer the final design?over designs rejected in the first 10 answers · bar: mean of 24 simulated people · dot: the person above0%50% chance100%Acquisition order (EUBO)at session end88%a week later78%Randomized orderat session end76%a week later73%
stable preferencethe person nowmodel's fitted utilityAcquisition order (EUBO) · one week later−0.6−0.4−0.20.00.20.40.6utilityearly favoritefinal designrejected earlyAnswers, first at topfirst 100.00.20.40.60.81.0design● chosen · ○ rejected · grey: no preference behind itRetest without the system:still prefers the final design?bar: mean of 24 people · dot: person above0%50% chance100%Acquisition (EUBO)session end88%week later78%Randomizedsession end76%week later73%
Figure 45.2 A simulation of a model, not data. The dashed curve is the simulated person's stable preference, the magenta curve their utility at the moment (stable part plus induced change), and the blue curve the fitted utility of a Gaussian process preference model (probit, Laplace), with the next pair chosen by EUBO (never repeating a pair) or at random. Curves are centered, because comparisons fix a utility only up to an added constant (Section 18.4). The middle panel lists the answers in order. The bottom panel retests, without the system, whether the final design (star) is still preferred to designs rejected in the first ten answers (crosses), at the end of the session and a week later, averaged over 24 simulated people; each runs under both query orders with the same random draws, which a real experiment cannot do. All parameter values are illustrative.

Things to try:

  • Play the default session. The blue curve follows the magenta bump that grows where the person keeps choosing, not the dashed curve, and the acquisition order keeps the current favorite in almost every query. The final design wins 88% of retests at the end of the session and 78% a week later under the acquisition order, against 76% and 73% under the randomized order.
  • Set Induced change to 0. Each order's two bars become equal (73% and 73% under the acquisition order, 71% and 71% under the randomized one). Now raise the noise to 1: agreement falls to 69% and 65%, the same at both times, since noise leaves no trace of time.
  • Return the noise to 0.3 and set incomplete answers to 0.4. Agreement falls toward chance (58% and 56%), again equally at both times. Tick Offer "can't compare": the model fits only real answers, and agreement recovers (74% and 67%).
  • Set the stable preference to 0 (other settings at their defaults). There is nothing to discover, yet the acquisition order produces a confident fit and an end-of-session agreement of 74%, falling to 59% a week later; under the randomized order, 55% and 52%.
  • Return the stable preference to 0.4 and set Lasting share to 1. Each order's bars are equal again (94% under the acquisition order, 85% under the randomized one): a retest cannot see change that lasts.

Two lessons follow, both statements about this model (inference). First, the posterior concentrates whenever the answers are consistent, whatever made them consistent, so concentration is not evidence that a preference has stabilized, and Section 46.6 does not use it as a stopping rule. Second, no single measurement separates the components. The clearest signature of induced change is how much agreement falls between the end of the session and a week later, and whether it falls more under the acquisition order; delayed agreement alone would mislead, since at the defaults the acquisition order's final designs still win more often a week later (78% against 73%), although by the stable preference they are no better.

45.4.2 What separates the components #

Table 45.4 states what each component does to the data and which measurement isolates it. The measurements are cheap additions to an ordinary session, and Section 47.4 turns them into one experiment.

Table 45.4 How each component shows up in the data, and the measurement that isolates it (inference).
Component Signature in the data Measurement that isolates it
Stable preference the same answers at any delay and under any query order delayed retest agreeing with the end-of-session retest
Structured evaluation noise lower consistency for hard or similar pairs, unrelated to the delay the same pair repeated at different lags; response times
Query-induced change answers that favor recently chosen designs; agreement that falls with delay, more under acquisition order; final designs close to the early favorite randomized query order compared with acquisition order, with a retest at the end of the session and a week later
Incomplete or randomized answers a false-preference rate on identical pairs; "can't compare" answers concentrated on complex pairs placebo pairs (the same design shown twice); a "can't compare" option
Sources cited in Section 45.4 9
  1. Shen et al. (2025b) Early versus late noise differentially enhances or degrades context-dependent choice
  2. Ghosal et al. (2023) The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
  3. Glickman and Sharot (2025) How human–AI feedback loops alter human perceptual, emotional and social judgements
  4. Vessel et al. (2018) Stronger shared taste for natural aesthetic domains than for artifacts of human culture
  5. Brielmann et al. (2024) Modelling individual aesthetic judgements over time
  6. Polanía et al. (2019) Efficient coding of subjective value
  7. Bavard and Palminteri (2023) The functional form of value normalization in human reinforcement learning
  8. Taddei et al. (2026) Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis
  9. Pombo et al. (2023) The intrinsic variance of beauty judgment

45.5 Estimate and intervention #

If a comparison mixes these components, the output of a PBO session is two things at once: an estimate of the person's preference and an intervention on it. A study that neither randomizes the order of queries nor retests after a delay cannot tell whether what converged was the preference or the system's effect on the person. Choosing is known to change preferences by a considerable amount, but no study has directly measured query-induced change in a PBO session; Section 47.4 describes the experiment that would. Its controls are the ones needed to judge whether a change the system caused was legitimate, so the methodological and the normative question can be answered in the same experiment (inference).

45.5.1 Why legitimacy conditions are needed #

Two results make legitimacy a technical question rather than an afterthought. When preferences can change, standards of legitimacy cannot be read off the objective function: each of the eight notions of alignment that Carroll et al. (2024) compare fails in one of two ways (Section 45.3.3). And learners optimized on user feedback learn to target the users who are easiest to influence (Williams et al., 2025). An acquisition function that rewards fast convergence or a clean signal could in principle favor queries that make the person more predictable over queries that serve their interests (inference), so a team should audit its own acquisition function and stopping rule for that incentive (Section 46.8).

45.5.2 Candidate conditions #

There is a set of operational candidates, though no consensus: respect the person's meta-preferences, their preferences about how their preferences may change (Ashton and Franklin, 2022), a workshop paper; keep influence overt, and give the system no incentive to make the user more predictable (Carroll et al., 2023); let the user choose the mode of autonomy (Fischli et al., 2026); require reflective endorsement with bounded influence that does not degrade the person's factual beliefs and keeps future options open while the preference is uncertain (Kanwal and Tran, 2026), a workshop paper; and judge a change from the standpoint of the selves both before and after it (Pettigrew, 2023).

Combining these, the defensible target of single-user PBO is not "the latent utility" but "the utility the user would reflectively endorse, within a protocol of bounded influence" (inference). Algorithm 46.2 makes that target operational: record the person's attitude toward having their taste changed, measure and bound the shift against a random or balanced query order, and retest the final design without the system's framing at the end of the session and again at the next one.

Sources cited in Section 45.5 7
  1. Carroll et al. (2024) AI Alignment with Changing and Influenceable Reward Functions
  2. Williams et al. (2025) On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
  3. Ashton and Franklin (2022) Solutions to preference manipulation in recommender systems require knowledge of meta-preferences
  4. Carroll et al. (2023) Characterizing Manipulation from AI Systems
  5. Fischli et al. (2026) Agents, Alignment, and the Many Faces of Autonomy
  6. Kanwal and Tran (2026) Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
  7. Pettigrew (2023) Nudging for changing selves

45.6 Settled, contested, missing #

Research status Settled, contested, missing

Settled. The acquisition rules of 2017 to 2021 have documented failures, and EUBO has a decision-theoretic justification. Regret theory for preference feedback offers only conditional upper bounds. Domain-general decision fatigue, ego depletion, induced-compliance dissonance, and moral licensing within a person did not survive large replications or bias correction. Choice-induced preference change, range normalization, efficient coding of value noise, and random-utility consistency for most people are well supported, and a sizable share of people report incomplete preferences when allowed.

Contested. Whether EUBO's collapse and the rank-deficient Hessian cost anything on human tasks (the evidence is from 2026 preprints). Whether complex valuation errors explain risky-choice anomalies. How much of the inconsistency in PBO sessions is preference formation rather than structured error around a stable target. Whether the hierarchical view, true of self-reported values, says anything about pairwise design judgments.

Missing. A direct measurement of query-induced change in a PBO session, with randomized query order and a delayed retest (Section 47.4). A test of the four-component decomposition in design tasks. A comparison of link functions on human data, and a model of drifting utility. A public data set of individual-level pairwise judgments with timestamps, presentation order, and repeated pairs. A demonstration, for or against, that acquisition functions favor queries that make people predictable. Agreed, tested conditions for when a change a system causes is legitimate.

45.7 Exercises #

Both exercises use Figure 45.2.

Exercise 45.1

Set the Lasting share slider to 1 and leave everything else at its default. The retest bars of each query order are now equal. Has query-induced change disappeared? What evidence, available in a real experiment, would still reveal it?

Solution

No. Induced change no longer fades, so the person's preference a week later is the changed one, and the retest has nothing to detect. The orders still differ: the acquisition order's final designs are preferred more strongly (94% against 85%) because its queries kept reinforcing one favorite, though by the stable preference they are no better. In a real experiment the remaining evidence is that people randomized to different query orders end with different distributions of final designs, those of the acquisition order lying close to each person's early favorite. Because an efficient acquisition rule also settles early, this evidence is weaker than a drop after a delay (inference). Whether such a lasting change is acceptable is a question of legitimacy (Section 45.5).

Exercise 45.2

Set the stable preference to 0 and induced change to 0.12. The fitted utility under the acquisition order has a clear peak, and the end-of-session retest agreement is high. Explain why the model is confident although the person has no stable preference, and what this implies for stopping rules.

Solution

The likelihood sees only which option won each comparison. Under the acquisition order the current favorite appears in most queries, and every win raises it further through induced change, so the answers are highly consistent, and consistency is all the posterior measures. A stopping rule that waits for the posterior to concentrate would stop here with confidence; a rule that also requires repeated pairs to agree across a delay, or the final design to survive a retest without the system, would not (Section 46.6).

Further reading #

References

  1. Agranov, M., and Ortoleva, P. (2025). Ranges of Randomization. Review of Economics and Statistics. Cited in §45.2
  2. Alós-Ferrer, C., Fehr, E., and Garagnani, M. (2023). Identifying Nontransitive Preferences. University of Zurich. working paper Cited in §45.2
  3. Ambuehl, S., Bernheim, B. D., and Lusardi, A. (2022). Evaluating Deliberative Competence: A Simple Method with an Application to Financial Choice. American Economic Review. Cited in §45.2
  4. Andersson, D., Lindberg, M., Tinghög, G., and Persson, E. (2025). No evidence for decision fatigue using large-scale field data from healthcare. Communications Psychology. Cited in §45.2
  5. Ashton, H., and Franklin, M. (2022). Solutions to preference manipulation in recommender systems require knowledge of meta-preferences. FAccTRec Workshop (RecSys 2022). workshop paper Cited in §45.5
  6. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §45.1
  7. Bagaïni, A., Liu, Y., Kapoor, M., Son, G., Bürkner, P.-C., Tisdall, L., and Mata, R. (2025). A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures. Nature Human Behaviour. doi:10.1038/s41562-024-02085-2. Cited in §45.1
  8. Bavard, S., and Palminteri, S. (2023). The functional form of value normalization in human reinforcement learning. eLife. Cited in §45.2 §45.4
  9. Bavard, S., Lebreton, M., Khamassi, M., Coricelli, G., and Palminteri, S. (2018). Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences. Nature Communications. Cited in §45.2
  10. Bernheim, B. D., and Rangel, A. (2009). Beyond Revealed Preference: Choice-Theoretic Foundations for Behavioral Welfare Economics *. Quarterly Journal of Economics. Cited in §45.2
  11. Bhatia, S., and Loomes, G. (2017). Noisy preferences in risky choice: A cautionary note. Psychological Review. Cited in §45.2
  12. Brielmann, A. A., Berentelg, M., and Dayan, P. (2024). Modelling individual aesthetic judgements over time. Philosophical Transactions of the Royal Society B: Biological Sciences. Cited in §45.3 §45.4
  13. Brown, A. L., Imai, T., Vieider, F. M., and Camerer, C. F. (2024). Meta-analysis of Empirical Estimates of Loss Aversion. Journal of Economic Literature. Cited in §45.2
  14. Carroll, M., Dragan, A., Russell, S., and Hadfield-Menell, D. (2022). Estimating and Penalizing Induced Preference Shifts in Recommender Systems. ICML 2022. Cited in §45.2
  15. Carroll, M., Chan, A., Ashton, H., and Krueger, D. (2023). Characterizing Manipulation from AI Systems. Equity and Access in Algorithms, Mechanisms, and Optimization. Cited in §45.5
  16. Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. Cited in §45.3 §45.5
  17. Cettolin, E., and Riedl, A. (2019). Revealed preferences under uncertainty: Incomplete preferences and preferences for randomization. Journal of Economic Theory. Cited in §45.2
  18. Chang, R. (2024). What’s so Hard about Hard Choices? Erasmus Journal for Philosophy and Economics. doi:10.23941/ejpe.v17i1.872. Cited in §45.3
  19. Chidambaram, K., Seetharaman, K. V., and Syrgkanis, V. (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Cited in §45.1
  20. Choung, O.-H., Vianello, R., Segler, M., Stiefl, N., and Jiménez-Luna, J. (2023). Extracting medicinal chemistry intuition via preference machine learning. Nature Communications. doi:10.1038/s41467-023-42242-1. Cited in §45.1
  21. Costa-Gomes, M. A., Cueva, C., Gerasimou, G., and Tejišcák, M. (2022). Choice, deferral, and consistency. Quantitative Economics. Cited in §45.3
  22. De Witte, S., Taets, J., Retzler, A., Crevecoeur, G., and Lefebvre, T. (2025). How to Capture Human Preference: Commissioning of a Robotic Use-Case via Preferential Bayesian Optimisation. arXiv. preprint Cited in §45.1
  23. Dean, S., and Morgenstern, J. (2022). Preference Dynamics Under Personalized Recommendations. EC 2022. Cited in §45.2
  24. Doumont, C., Fan, D., Maus, N., Gardner, J. R., Moss, H., and Pleiss, G. (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). Cited in §45.1
  25. Engelmann, D., and Hollard, G. (2010). Reconsidering the Effect of Market Experience on the "Endowment Effect". Econometrica. Cited in §45.3
  26. Enisman, M., Shpitzer, H., and Kleiman, T. (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Cited in §45.2
  27. Evren, Ö., and Ok, E. A. (2011). On the multi-utility representation of preference relations. Journal of Mathematical Economics. Cited in §45.2
  28. Fischli, R., Franklin, M., Manzini, A., and Gabriel, I. (2026). Agents, Alignment, and the Many Faces of Autonomy. Minds and Machines. Cited in §45.5
  29. Fosgerau, M., Melo, E., de Palma, A., and Shum, M. (2020). Discrete Choice and Rational Inattention: A General Equivalence Result. International Economic Review. Cited in §45.3
  30. Fraser, I., Balcombe, K., Williams, L., and McSorley, E. (2021). Preference stability in discrete choice experiments. Some evidence using eye-tracking. Journal of Behavioral and Experimental Economics. Cited in §45.3
  31. Frederick, S., Lee, L., and Baskin, E. (2014). The Limits of Attraction. Journal of Marketing Research. Cited in §45.2
  32. Frömer, R., Dean Wolf, C. K., and Shenhav, A. (2019). Goal congruency dominates reward value in accounting for behavioral and neural correlates of value-based decision-making. Nature Communications. Cited in §45.3
  33. Ge, L., Halpern, D., Micha, E., Procaccia, A., Shapira, I., Vorobeychik, Y., and Wu, J. (2024). Axioms for AI Alignment from Human Feedback. Advances in Neural Information Processing Systems 37. Cited in §45.2
  34. Ghosal, G. R., Zurek, M., Brown, D. S., and Dragan, A. D. (2023). The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types. AAAI. Cited in §45.4
  35. Glickman, M., and Sharot, T. (2025). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour. doi:10.1038/s41562-024-02077-2. Cited in §45.4
  36. Hardt, M., and Mendler-Dünner, C. (2025). Performative Prediction: Past and Future. Statistical Science. Cited in §45.2
  37. Heckel, R., Shah, N. B., Ramchandran, K., and Wainwright, M. J. (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. Cited in §45.1
  38. Hendrickx, J. M., Olshevsky, A., and Saligrama, V. (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. Cited in §45.2
  39. Jiang, X., Lim, L.-H., Yao, Y., and Ye, Y. (2011). Statistical ranking and combinatorial Hodge theory. Mathematical Programming. Cited in §45.2
  40. Kanwal, M., and Tran, C. (2026). Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction. AAAI-26 Workshop on Machine Ethics. workshop paper Cited in §45.5
  41. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §45.1
  42. Keswani, V., Cousins, C., Nguyen, B., Conitzer, V., Heidari, H., Borg, J. S., and Sinnott-Armstrong, W. (2026). Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback. AAAI. Cited in §45.1
  43. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §45.1
  44. Li, S., Zhang, Y., Ren, Z., Liang, C., Li, N., and Shah, J. A. (2024a). Enhancing Preference-based Linear Bandits via Human Response Time. Advances in Neural Information Processing Systems. Cited in §45.1
  45. Li, Y., Parashar, A., Zhou, E., and Fan, C. (2026c). Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference. arXiv preprint 2602.06029. preprint Cited in §45.3
  46. McCausland, W. J., Davis-Stober, C., Marley, A., Park, S., and Brown, N. (2020). Testing the Random Utility Hypothesis Directly. The Economic Journal. doi:10.1093/ej/uez039. Cited in §45.2 §45.3
  47. Millidge, B., Tschantz, A., and Buckley, C. L. (2021). Whence the Expected Free Energy? Neural Computation. Cited in §45.3
  48. Nielsen, K., and Rehbeck, J. (2022). When Choices Are Mistakes. American Economic Review. Cited in §45.3
  49. Nielsen, K., and Rigotti, L. (2026). Revealed Incomplete Preferences. Working paper (author's website). working paper Cited in §45.2
  50. Nieuwenstein, M. R., Wierenga, T., Morey, R. D., Wicherts, J. M., Blom, T. N., Wagenmakers, E.-J., and van Rijn, H. (2015). On making the right choice: A meta-analysis and large-scale replication attempt of the unconscious thought advantage. Judgment and Decision Making. Cited in §45.3
  51. Niwa, R., Yoshida, S., Koyama, Y., and Ushiku, Y. (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Cited in §45.3
  52. Oh, G., Lee, J., Park, J., Yu, Y., Bae, W., and Noh, J. (2026b). Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs. ICLR 2026 Workshop: I Can't Believe It's Not Better (ICBINB). workshop paper Cited in §45.1
  53. Ok, E. A., and Tserenjigmid, G. (2022). Indifference, indecisiveness, experimentation, and stochastic choice. Theoretical Economics. Cited in §45.2
  54. Oprea, R. (2024). Decisions under Risk Are Decisions under Complexity. American Economic Review. Cited in §45.3
  55. Ou, C., Buschek, D., Mayer, S., and Butz, A. (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §45.3
  56. Pásztor, B., Kassraie, P., and Krause, A. (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Cited in §45.1
  57. Peng, Y.-H., Bigham, J. P., and Wu, J. (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Cited in §45.1
  58. Perdomo, J. C., Zrnic, T., Mendler-Dünner, C., and Hardt, M. (2020). Performative Prediction. ICML. Cited in §45.2
  59. Pettigrew, R. (2023). Nudging for changing selves. Synthese. Cited in §45.3 §45.5
  60. Polanía, R., Woodford, M., and Ruff, C. C. (2019). Efficient coding of subjective value. Nature Neuroscience. Cited in §45.2 §45.4
  61. Pombo, M., Brielmann, A. A., and Pelli, D. G. (2023). The intrinsic variance of beauty judgment. Attention, Perception, & Psychophysics. Cited in §45.4
  62. Prat-Carrabin, A., and Woodford, M. (2022). Efficient coding of numbers explains decision bias and noise. Nature Human Behaviour. Cited in §45.2
  63. Rotella, A., Jung, J., Chinn, C., and Barclay, P. (2026). Observation Moderates the Moral Licensing Effect: A Meta-Analytic Test of Interpersonal and Intrapsychic Mechanisms. Personality and Social Psychology Bulletin. Cited in §45.2
  64. Schäfer, N., Zhao, G., Li, B., Kupnik, M., Seyfarth, A., Beckerle, P., and Grimmer, M. (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §45.1
  65. Schoinas, E., Rastogi, A., Carter, A., Granley, J., and Beyeler, M. (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Cited in §45.1
  66. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. J. (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Cited in §45.2
  67. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §45.1
  68. Shen, B., Nguyen, D., Wilson, J., Glimcher, P. W., and Louie, K. (2025b). Early versus late noise differentially enhances or degrades context-dependent choice. Nature Communications. doi:10.1038/s41467-025-59140-3. Cited in §45.4
  69. Shevlin, B. R. K., Smith, S. M., Hausfeld, J., and Krajbich, I. (2022). High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity. Proceedings of the National Academy of Sciences. Cited in §45.2
  70. Shvartsman, M., Letham, B., Bakshy, E., and Keeley, S. (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Cited in §45.1
  71. Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Cited in §45.1
  72. Stango, V., and Zinman, J. (2024). Behavioral Biases Are Temporally Stable. Working paper (author's website). working paper Cited in §45.3
  73. Strang, A., Abbott, K. C., and Thomas, P. J. (2022). The Network HHD: Quantifying Cyclic Competition in Trait-Performance Models of Tournaments. SIAM Review. Cited in §45.2
  74. Taddei, S., Koppen, W., Alfio, E., Nuzzo, S., Flynn, L., Diaz, M. A., … Verstraten, T. (2026). Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis. arXiv. preprint Cited in §45.4
  75. Vaidis, D. C., Sleegers, W. W. A., van Leeuwen, F., DeMarree, K. G., Sætrevik, B., Ross, R. M., … Priolo, D. (2024). A Multilab Replication of the Induced-Compliance Paradigm of Cognitive Dissonance. Advances in Methods and Practices in Psychological Science. Cited in §45.2
  76. Vecchione, M., Schwartz, S. H., Davidov, E., Cieciuch, J., Alessandri, G., and Marsicano, G. (2020). Stability and change of basic personal values in early adolescence: A 2‐year longitudinal study. Journal of Personality. Cited in §45.3
  77. Vessel, E. A., Maurer, N., Denker, A. H., and Starr, G. G. (2018). Stronger shared taste for natural aesthetic domains than for artifacts of human culture. Cognition. Cited in §45.4
  78. Viappiani, P., and Boutilier, C. (2020). On the equivalence of optimal recommendation sets and myopically optimal query sets. Artificial Intelligence. Cited in §45.1
  79. Vohs, K. D., Schmeichel, B. J., Lohmann, S., Gronau, Q. F., Finley, A. J., Ainsworth, S. E., … Albarracín, D. (2021). A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect. Psychological Science. Cited in §45.2
  80. Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., and Dragan, A. (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Cited in §45.5
  81. Wu, K., and Gardner, J. R. (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Cited in §45.1
  82. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §45.1
  83. Yap, S. C. Y., Wortman, J., Anusic, I., Baker, S. G., Scherer, L. D., Donnellan, M. B., and Lucas, R. E. (2017). The effect of mood on judgments of subjective well-being: Nine tests of the judgment model. Journal of Personality and Social Psychology. Cited in §45.2
  84. Yechiam, E., and Zeif, D. (2025). Loss aversion is not robust: A re-meta-analysis. Journal of Economic Psychology. Cited in §45.2
  85. Zhang, X., Huang, D., Kaski, S., and Martinelli, J. (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §45.1
  86. Zhang, R., Zhu, X., Pourebadi Khotbehsara, M., Dao, W., Bıyık, E., and Culbertson, H. (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Cited in §45.1
  87. Zylberberg, A., Bakkour, A., Shohamy, D., and Shadlen, M. N. (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Cited in §45.2