What a Comparison Measures
Preferential Bayesian optimization (PBO) looks for the design a person likes best by asking them to compare options. A statistical model reads each answer as a noisy comparison of an unobserved number attached to every design, the person's utility: by default a Gaussian process prior (a distribution over smooth functions) on the utility function, and a probit link that turns the difference between two utilities into a choice probability (Chapter 16, Chapter 18). An acquisition function picks the next pair (Chapter 19), and after a few dozen to a couple of hundred answers the design with the highest estimated utility is the result.
That model assumes one fixed utility function and noise of constant size. This chapter gathers what the earlier parts found about that assumption to answer a plain question: what does one answer to "which do you prefer?" measure? Our answer is four things at once, so a PBO session is both an estimate of a preference and an intervention on it. A sentence that is our own inference ends with "(inference)".
45.1 The bottleneck moved #
Between 2017 and 2026, most of the field's effort went into acquisition functions and regret theory, and both advanced. The bottleneck has since moved from algorithms to measurement: what a single comparison measures, how answers should be modeled, and what asking does to the person who answers. The largest errors now sit in the observation model, the part of the model that says how a person's answer arises from their utility, and in how methods are evaluated against real people.
45.1.1 Acquisition: from heuristics to decision theory #
The acquisition rules common from 2017 to 2021 adapted scalar Bayesian optimization (Section 28.1); four independent groups reported that one of them, modified expected improvement, stalls (Section 28.4), and on a constructed instance batch expected improvement is not asymptotically consistent: even unlimited queries need not find the best option (Astudillo et al., 2023). The rule that replaced these heuristics is EUBO, the expected utility of the best option, which scores a pair by the expected utility of whichever member the person would pick (Section 19.4); qEUBO is its multi-option form. EUBO is one-step Bayes optimal, the best possible choice if only one question remained, a property with an older root in decision analysis: the optimal set of options to recommend is also the myopically optimal set to ask about (Viappiani and Boutilier, 2020). Its flaws were documented in 2026, in two preprints: its queries collapse toward the estimated maximum, and its pairs tend to form isolated pieces of the comparison graph, which leaves the Hessian of the Laplace approximation rank deficient (Wu and Gardner, 2026; Shao et al., 2026). Section 19.6 tells both stories, and Section 19.5.1 shows a similar collapse in a recorded simulation.
Rankings of acquisition functions flip with the metric, the noise model, and
the inference method: in the paper that introduced the optimistic algorithm
POP-BO, qEUBO reported a slightly better final solution while its cumulative
regret, the summed shortfall of the options shown along the way (Chapter 13),
was more than 2.5 times higher (Xu et al., 2024b). qEUBO with BoTorch's
PairwiseGP model is the best-maintained default implementation
(Section 27.5), a standing that comes from the software ecosystem, not
from comparisons with simpler methods.
45.1.2 Theory: mature, conditional, and only upper bounds #
Regret theory for preference feedback matured, but every result carries conditions, and all are upper bounds: guarantees that regret grows no faster than some rate in the number of queries . Two quantities recur. , the maximum information gain of the kernel, measures how much noisy observations can reveal about a function drawn from the Gaussian process prior (Section 6.5). is an upper bound on the inverse slope of the link function, which becomes large when choice probabilities saturate near 0 or 1 (Section 29.5). The sequential algorithms MaxMinLCB and PF-TS have preference-probability regret , where ignores logarithmic factors, with constants that grow with (Pásztor et al., 2024; Lazzaro et al., 2026), and a batched algorithm does better only under extra conditions (Kayal et al., 2025) (Section 29.4). There is no kernelized lower bound under a logistic or probit link (Section 29.7). And the theory covers elimination, optimistic, and Thompson algorithms on frequentist kernel estimators, while practice runs a Laplace posterior with EUBO, for which there are only one-step optimality and finite-domain consistency results; nothing connects the two (inference).
Whether a comparison costs more than a number depends on the goal. For the regret of a single fixed utility, the upper bounds in the finite-arm, linear, and kernel settings show that pairwise comparisons are not more expensive in order. For identifying a heterogeneous population, pairwise data are the weakest feedback: a model fitted to binary comparisons implicitly aggregates by Borda count, which scores each option by its average chance of beating the others (Siththaranjan et al., 2024), and rankings of at least three options are needed to identify latent user types (Chidambaram et al., 2026).
45.1.3 The observation model is the main bottleneck #
The probit link with constant noise remains the default, and its extensions mostly come from one research group with one evaluation each (Chapter 27). Human answers depart a long way from "a fixed utility plus noise of constant size". When 20 people with design training judged the same 600 pairs of generated interfaces, their agreement was 0.25 on Krippendorff's , a chance-corrected agreement score on which 0 is chance and 1 is perfect agreement, in a 2026 preprint (Peng et al., 2026). Among 35 chemists, Fleiss' kappa, a score of the same kind, was 0.40 and 0.32 in two rounds (Choung et al., 2023). The same moral pairwise questions, repeated within and across sessions, had on average 6% to 20% of answers flip (Keswani et al., 2026). In retinal-implant optimization, sighted participants and the simulated agents meant to stand in for them agreed on only about 50% of choices (Schoinas et al., 2025). And in a meta-analysis of the stability of risk preference, the estimated reliability (the correlation between two measurements of the same people taken close together) was 0.61 for self-reported propensity to take risks and only 0.25 for behavioral measures, standardized tasks such as choices between lotteries (Bagaïni et al., 2025). A pairwise choice is a behavioral measure, so a few stated-preference questions alongside the comparisons might carry the stable component (inference). Figure 45.1 puts these numbers side by side.
Some things to try:
- Choose the flips row. A flip rate of 6% to 20% means the same person repeats 80% to 94% of their answers; a person answering by coin flip would repeat 50%.
- Compare the two reliability rows. They come from one meta-analysis and use one statistic, so they can be compared: self-reports (0.61) are more than twice as reliable as behavioral measures (0.25).
Simulation studies of acquisition functions assume an observation model, so they cannot measure the error that comes from that model being wrong; improving the likelihood is more likely to change research conclusions than another acquisition function (inference). Of the signals that cost the person no extra effort, response times (Shvartsman et al., 2024; Li et al., 2024a) and stated confidence (Zhang et al., 2026b) have improved models on real human data, so far in offline fits rather than in closed-loop optimization.
45.1.4 Small samples, weak evaluation #
Studies of pure preference feedback in interactive design mostly enroll 6 to 60 people, applied studies elsewhere 1 to 35 evaluators with a median below 10, and validation is almost always internal: the same person later chooses the result (Chapter 32, Section 34.4). Benchmarks are weak in the same direction: scalar test functions in 1 to 8 dimensions, at least six noise models and five definitions of regret, no shared benchmark, and no public data set of individual-level pairwise judgments (Chapter 31). Even an expert's own cost function can miss their choices: in a robot-commissioning study with a single expert operator, reported in a 2025 preprint, a cost function the expert had designed did not fully capture the expert's own choices, even after its weights were refitted to the expert's highest-rated experiments (De Witte et al., 2025). Synthetic decision makers defined by a hidden cost function may therefore overstate how well methods work on real people (inference).
45.1.5 Simpler methods often do as well #
In several settings, simpler methods perform comparably: people tuning an exoskeleton themselves reached, in about 11 minutes, metabolic savings of the same order as algorithmic tuning (Schäfer et al., 2026); a spherical input mapping with Bayesian linear regression reached the state of the art on high-dimensional benchmarks (Doumont et al., 2026); and random selection was hard to beat in online direct preference optimization of language models, in a 2026 workshop paper (Oh et al., 2026b), while the authors of the amortized optimizer PABBO write that a random strategy often beats some Gaussian process baselines (Zhang et al., 2025a). Theory points to where PBO's advantage must come from: when ranking items actively from noisy comparisons, parametric assumptions such as Bradley-Terry or Thurstone buy at most a logarithmic gain (Heckel et al., 2019), so PBO's sample efficiency should come mainly from the kernel sharing information between neighboring designs, not from the link function (inference).
The setting in which PBO remains supported by evidence is therefore narrow: options that can be judged only by perceptual comparison, no numerical objective, a dimension already low after the representation has been designed, and a budget of tens to about two hundred comparisons. Section 46.9 turns this into advice, including the baselines a study needs before it can claim an advantage for PBO.
Acquisition functions now have a decision-theoretic footing and regret bounds have caught up with scalar feedback, but the error that matters most comes from what the model assumes about the person, which simulations with a known observation model cannot see (inference).
Sources cited in Section 45.1 24
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Viappiani and Boutilier (2020) On the equivalence of optimal recommendation sets and myopically optimal query sets
- Wu and Gardner (2026) Knowledge Gradient for Preference Learning
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Pásztor et al. (2024) Bandits with Preference Feedback: A Stackelberg Game Perspective
- Lazzaro et al. (2026) A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
- Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
- Siththaranjan et al. (2024) Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
- Chidambaram et al. (2026) Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
- Peng et al. (2026) Efficient Personalization of Generative User Interfaces
- Choung et al. (2023) Extracting medicinal chemistry intuition via preference machine learning
- Keswani et al. (2026) Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
- Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
- Bagaïni et al. (2025) A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures
- Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences
- Li et al. (2024a) Enhancing Preference-based Linear Bandits via Human Response Time
- Zhang et al. (2026b) Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
- De Witte et al. (2025) How to Capture Human Preference: Commissioning of a Robotic Use-Case via Preferential Bayesian Optimisation
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Doumont et al. (2026) We Still Don't Understand High-Dimensional Bayesian Optimization
- Oh et al. (2026b) Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
- Zhang et al. (2025a) PABBO: Preferential Amortized Black-Box Optimization
- Heckel et al. (2019) Active ranking from pairwise comparisons and when parametric assumptions do not help
45.2 What the disciplines changed #
Part IX collected results about preference from psychology, economics, neuroscience, and other fields. Since 2017, some failed in preregistered replications (studies whose hypotheses and analyses are fixed before data collection) or shrank in meta-analyses corrected for publication bias; others were strengthened and gained mechanisms that can be written into a likelihood. Table 45.1 and Table 45.2 list only the results with direct consequences for preferential optimization; the last column of each is our inference. The effect sizes and express a difference between conditions in standard deviations, so is a twentieth of a standard deviation and is a moderate effect.
| Result | Status | Key evidence | What it means for PBO (inference) |
|---|---|---|---|
| Ego depletion and domain-general decision fatigue: deciding draws on a limited resource | weakened to near zero | 36 laboratories, (Vohs et al., 2021); 231,076 triage calls show no decision fatigue (Andersson et al., 2025) | Do not set session length on this basis; measure fatigue within the task |
| Induced-compliance dissonance: people change attitudes to match what they were induced to do | not replicated | 39 laboratories, (Vaidis et al., 2024) | Explain post-choice preference change by revaluation, not by classic dissonance |
| Moral licensing within a person: a good deed licenses a later lapse | near zero after bias correction; only a reputational effect when observed remains | between and (Rotella et al., 2026) | Do not model compensation between successive answers in private sessions |
| Mood as information: current mood is read as evidence about the options | much weakened | most of nine replications not significant (Yap et al., 2017) | Do not schedule exploration by mood |
| Value discrimination worsens near the optimum | fails for value judgments; holds for perceptual parameters | decisions between high-value options are faster and more accurate (Shevlin et al., 2022) | Value comparisons may become more accurate near convergence; perceptual parameters keep a just-noticeable-difference floor |
| A loss-aversion coefficient of about 2.25 | value unstable | meta-analytic mean 1.955 (Brown et al., 2024); about 1.07 when symmetric and unsorted (Yechiam and Zeif, 2025) | Do not fix an asymmetric likelihood around the current best |
| The attraction (decoy) effect: adding a dominated option boosts its neighbor | limited to attributes shown as numbers | Frederick et al. (2014) | The mechanism is weak in perceptual design |
| Result | Status | Key evidence | What it means for PBO (inference) |
|---|---|---|---|
| Choice-induced preference change: choosing an option raises its value for the chooser | strengthened | (Enisman et al., 2021); each choice adds about $0.18 to the chosen item (Zylberberg et al., 2024) | Add a choice-history term to the likelihood; the most-compared current best is affected most |
| Range normalization: values are coded relative to the range on offer | strengthened | Bavard et al. (2018); Bavard and Palminteri (2023) | The utility's scale is session-relative; recalibrate before reusing it across sessions |
| Efficient coding of value noise: precision goes to the values a person expects to see | new and strengthened | Polanía et al. (2019); Prat-Carrabin and Woodford (2022) | Noise scales with the values presented, so fixed noise misreads late-session precision |
| Random utility consistency: repeated choices behave as if drawn from a distribution over utilities | strengthened for most people; part of the observed intransitivity is genuine | most of 141 participants satisfy random utility (McCausland et al., 2020); with response times separating noise from preference, 19.24% and 39.58% of the transitivity violations in two data sets are genuine, and the rest could be either (Alós-Ferrer et al., 2023); the noise specification changes the inference (Bhatia and Loomes, 2017) | A scalar utility plus noise works for most people, but the link and noise form are substantive modeling choices |
| Incomplete preferences and deliberate randomization | strengthened, with limits | 40% to 50% of participants report incompleteness when allowed, on average 3.3 of 50 comparisons (Nielsen and Rigotti, 2026); about half choose inconsistently with complete preferences plus certainty independence, an axiom about mixing options with a sure outcome (Cettolin and Riedl, 2019); most choose to randomize (Agranov and Ortoleva, 2025) | Offer a "can't compare" answer and model it separately from ties and noise |
| Information per comparison | at most one bit, but the error rate is of the same order as with numerical ratings | Shah et al. (2016) | The structure of the comparison graph, not the bits per query, governs the error |
Two of these studies are working papers (Alós-Ferrer et al., 2023; Nielsen and Rigotti, 2026). The rows are developed in Part IX: fatigue in Section 37.2.1, moral licensing in Section 38.5, mood in Section 38.4, loss aversion in Section 40.1, context effects in Section 37.2.3, choice-induced change in Section 37.2.2, range normalization in Section 37.3.1, efficient coding in Section 39.4, random utility in Section 37.4.1, and incomplete preferences in Section 40.4 and Section 43.4.2.
The changes follow a pattern (inference). What weakened were mostly arguments for session-design rules (set session length by decision fatigue, schedule queries by mood) and for fixed numbers for directional biases (a loss-aversion coefficient, compensation in moral licensing). What strengthened were mechanisms with a magnitude that map onto parts of a likelihood. The evidence therefore supports a few mechanisms whose sign and size are estimated in each task, not one likelihood term for every documented bias.
Formal tools PBO has not yet used. Five families of results can be applied to PBO, and as of September 2026 we found no PBO paper that uses them:
- comparison-graph theory, in which the comparison graph's Laplacian spectrum and effective resistances govern the error of pairwise estimates (Shah et al., 2016; Hendrickx et al., 2019), and the Hodge decomposition splits comparison data into a part a scalar utility can explain and a cyclic part none can (Jiang et al., 2011; Strang et al., 2022) (Section 43.3.1, Section 43.3.2);
- performative prediction, which separates learning a target from steering it (Perdomo et al., 2020; Hardt and Mendler-Dünner, 2025) (Section 42.2);
- trust regions on induced preference shift (Carroll et al., 2022), needed because when preferences move toward what a system shows, low regret can be achieved trivially by moving them (Dean and Morgenstern, 2022);
- behavioral welfare economics, which calls one option better for a person's welfare only if the other is never chosen over it in any frame (Bernheim and Rangel, 2009) and measures welfare losses from misunderstood consequences (Ambuehl et al., 2022);
- multi-utility representations of incomplete preferences, in which an option counts as better only if every utility in a set agrees (Evren and Ok, 2011), with identification results that distinguish indifference, indecisiveness, and experimentation (Ok and Tserenjigmid, 2022).
The first and last apply to PBO logs and posteriors without new theory, so they are the place to start (inference). Relatedly, fitting one Bradley-Terry utility to several people's pooled comparisons is itself an aggregation rule, and it violates Pareto optimality and pairwise majority consistency (Ge et al., 2024).
Sources cited in Section 45.2 34
- Vohs et al. (2021) A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect
- Andersson et al. (2025) No evidence for decision fatigue using large-scale field data from healthcare
- Vaidis et al. (2024) A Multilab Replication of the Induced-Compliance Paradigm of Cognitive Dissonance
- Rotella et al. (2026) Observation Moderates the Moral Licensing Effect: A Meta-Analytic Test of Interpersonal and Intrapsychic Mechanisms
- Yap et al. (2017) The effect of mood on judgments of subjective well-being: Nine tests of the judgment model
- Shevlin et al. (2022) High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity
- Brown et al. (2024) Meta-analysis of Empirical Estimates of Loss Aversion
- Yechiam and Zeif (2025) Loss aversion is not robust: A re-meta-analysis
- Frederick et al. (2014) The Limits of Attraction
- Enisman et al. (2021) Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm
- Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
- Bavard et al. (2018) Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences
- Bavard and Palminteri (2023) The functional form of value normalization in human reinforcement learning
- Polanía et al. (2019) Efficient coding of subjective value
- Prat-Carrabin and Woodford (2022) Efficient coding of numbers explains decision bias and noise
- McCausland et al. (2020) Testing the Random Utility Hypothesis Directly
- Alós-Ferrer et al. (2023) Identifying Nontransitive Preferences
- Bhatia and Loomes (2017) Noisy preferences in risky choice: A cautionary note
- Nielsen and Rigotti (2026) Revealed Incomplete Preferences
- Cettolin and Riedl (2019) Revealed preferences under uncertainty: Incomplete preferences and preferences for randomization
- Agranov and Ortoleva (2025) Ranges of Randomization
- Shah et al. (2016) Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
- Hendrickx et al. (2019) Graph Resistance and Learning from Pairwise Comparisons
- Jiang et al. (2011) Statistical ranking and combinatorial Hodge theory
- Strang et al. (2022) The Network HHD: Quantifying Cyclic Competition in Trait-Performance Models of Tournaments
- Perdomo et al. (2020) Performative Prediction
- Hardt and Mendler-Dünner (2025) Performative Prediction: Past and Future
- Carroll et al. (2022) Estimating and Penalizing Induced Preference Shifts in Recommender Systems
- Dean and Morgenstern (2022) Preference Dynamics Under Personalized Recommendations
- Bernheim and Rangel (2009) Beyond Revealed Preference: Choice-Theoretic Foundations for Behavioral Welfare Economics *
- Ambuehl et al. (2022) Evaluating Deliberative Competence: A Simple Method with an Application to Financial Choice
- Evren and Ok (2011) On the multi-utility representation of preference relations
- Ok and Tserenjigmid (2022) Indifference, indecisiveness, experimentation, and stochastic choice
- Ge et al. (2024) Axioms for AI Alignment from Human Feedback
45.3 Found or made? #
Behind these results lies an old question: do repeated comparisons find a preference that was there before, or do they partly make it? The constructive view holds that preferences are built during elicitation, so the measurement changes what it measures; the stable view holds that there is an underlying preference that careful measurement can recover. Preferential optimization takes a side whether its users notice or not. The evidence since 2017 leaves part of each view standing and supports neither extreme.
45.3.1 What each view keeps #
The constructive view keeps its central claim: elicitation changes the object it measures, by a considerable amount and through mechanisms that can be modeled. Choosing an option raises its value (Table 45.2); for 59 people rating 55 morphed dog images, individual models with a learning component explained on average 17% more variance for the real presentation order than for a simulated random order (Brielmann et al., 2024); and forced choice manufactures some inconsistency that would otherwise show up as hesitation (Costa-Gomes et al., 2022). It loses some supporting arguments: the advantage of unconscious thought was not found in a meta-analysis and large replication (Nieuwenstein et al., 2015), the induced-compliance and mood effects failed or shrank (Table 45.1), and Ruth Chang's work on hard choices, sometimes read as calling a forced choice between options "on a par" a category error, treats such a choice as an occasion for commitment (Chang, 2024).
The stable view keeps the claim that there is a component worth estimating. Most people's repeated choices satisfy random utility (McCausland et al., 2020); behavioral biases are nearly unchanged at the population level over three years, in a working paper (Stango and Zinman, 2024); and some anomalies in risky choice look like errors in valuing complex options rather than preferences (Oprea, 2024), though that conclusion is disputed. It loses its strongest inference, that iterated querying converges to the true underlying preference. Plott's "discovered preference" hypothesis has been observed only after forced trading or after people reflected on axioms they themselves endorse (Nielsen and Rehbeck, 2022; Engelmann and Hollard, 2010), not as a natural outcome of repeated comparison, and in the field PBO loops often end before they converge, most evaluation sequences of a three-month deployment stopping at the first iteration (Ou et al., 2022) (Section 46.6).
45.3.2 The hierarchical view #
One position in this debate, which we call the hierarchical view, holds that preferences over ultimate goals (comfort, safety, beauty) always exist, while preferences over intermediate outcomes, such as a particular parameter setting, may have to be elicited. As a description of where stability lives, the evidence largely supports it: stability concentrates at the broad, abstract, self-reported level, and in a longitudinal study in early adolescence, 75% of respondents had value hierarchies that correlated at least 0.85 across two years (Vecchione et al., 2020).
As a prediction about PBO, the view fares worse. The value computed during a choice is relative to the current goal: goal congruence explains choices and their neural correlates better than reward value (Frömer et al., 2019), so a stable ultimate goal need not become a stable utility over pairwise comparisons (inference). Iterated querying has not been seen to converge (Section 45.3.1), and because stability is measured with self-reports and construction with choices, Section 47.1 describes the study that would test both levels on one design task. Nor does rational inattention, the theory that people spend costly attention optimally, single out the logit link as the view's defenders sometimes claim: it yields the logit only under a Shannon-entropy cost of information, and it can generate any additive random utility model (Fosgerau et al., 2020).
45.3.3 Preference scaffolding #
A second position, preference scaffolding, holds that PBO is not a readout of a fixed utility function but a structure that helps a person form a preference, so that evaluation should cover the exploration process and the person's reflective endorsement of the result, not only the final design. The core holds up, with three qualifications. Not all inconsistency is preference formation: in repeated discrete choice tasks, instability concentrates on options close in utility and on hard tasks, and the underlying preferences of unstable respondents do not differ from those of stable ones (Fraser et al., 2021). Scaffolding is not neutral: when preferences can be influenced, each of the eight notions of alignment compared in one analysis either errs toward undesirable influence or is overly risk-averse (Carroll et al., 2024), so helping and steering have to be told apart by conditions the objective does not supply (Section 45.5). And endorsement after the fact cannot by itself justify a change the system caused, because the changed person may endorse the result only because their values were changed (Pettigrew, 2023).
One argument offered for scaffolding does not hold. Active inference, a theory in which action minimizes surprise relative to preferred observations, is sometimes said to dissolve the dispute between construction and revelation, but it only re-encodes it: in Bayesian optimization it reduces to weighted information-gain acquisition functions whose weights still need tuning (Millidge et al., 2021; Li et al., 2026c) (Section 39.5). The most direct support is a 2025 measurement: when Bayesian optimization led the search, designers reached better results but reported significantly less sense of agency, and collaboration through explicit constraints matched collaboration through natural language in performance while giving more agency (Niwa et al., 2025). That supports a scaffold led by the designer, not preference shaping led by the system.
Sources cited in Section 45.3 19
- Brielmann et al. (2024) Modelling individual aesthetic judgements over time
- Costa-Gomes et al. (2022) Choice, deferral, and consistency
- Nieuwenstein et al. (2015) On making the right choice: A meta-analysis and large-scale replication attempt of the unconscious thought advantage
- Chang (2024) What’s so Hard about Hard Choices?
- McCausland et al. (2020) Testing the Random Utility Hypothesis Directly
- Stango and Zinman (2024) Behavioral Biases Are Temporally Stable
- Oprea (2024) Decisions under Risk Are Decisions under Complexity
- Nielsen and Rehbeck (2022) When Choices Are Mistakes
- Engelmann and Hollard (2010) Reconsidering the Effect of Market Experience on the "Endowment Effect"
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Vecchione et al. (2020) Stability and change of basic personal values in early adolescence: A 2‐year longitudinal study
- Frömer et al. (2019) Goal congruency dominates reward value in accounting for behavioral and neural correlates of value-based decision-making
- Fosgerau et al. (2020) Discrete Choice and Rational Inattention: A General Equivalence Result
- Fraser et al. (2021) Preference stability in discrete choice experiments. Some evidence using eye-tracking
- Carroll et al. (2024) AI Alignment with Changing and Influenceable Reward Functions
- Pettigrew (2023) Nudging for changing selves
- Millidge et al. (2021) Whence the Expected Free Energy?
- Li et al. (2026c) Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
45.4 Four components of a comparison #
Putting the pieces together, we propose that a single pairwise answer mixes four components, whose proportions are moderated by two further factors (Table 45.3; inference). No single paper has tested this decomposition, and the proportions change with familiarity with the domain, the similarity of the options, and the length of the session, so they have to be estimated in each application rather than assumed.
| Component | What it is | Evidence, and where it comes from | Representation in the model |
|---|---|---|---|
| Stable preference | an estimable latent utility | tests of random utility, stability of biases, value hierarchies; from lotteries, money choices, and self-reports, not from design comparisons | the latent utility of a Gaussian process or other surrogate |
| Structured evaluation noise | noise that varies with difficulty, similarity, protocol, ambiguity of the representation, and the distribution of values presented | early versus late noise (Shen et al., 2025b); efficient coding; range normalization; mechanisms supported by experiments, only offline fits in PBO | heteroscedastic noise; a joint likelihood with response times; separate noise scales per feedback type (Ghosal et al., 2023) |
| Query-induced change | revaluation after choosing, anchoring on the system's proposals, familiarity and fatigue | choice-induced preference change; biases grow after interacting with a biased AI (Glickman and Sharot, 2025); magnitudes from free-choice paradigms, never measured in design comparisons | a choice-history term; a drift term; balanced query order as a control |
| Incomplete or deliberately randomized answers | the person cannot or will not compare, or deliberately randomizes | experiments on incomplete and randomized preferences; from lotteries and money choices; the share in design comparisons is unknown | a "can't compare" answer; a mixture parameter; a multi-utility representation |
| Moderator 1: how much taste is shared | the informativeness of a population prior varies by domain | shared taste is high for faces and landscapes and low for architecture and artworks; where landscapes and exterior architecture were compared within the same people, the judgments were equally reliable (Vessel et al., 2018); individual models against the group average, median correlation 0.65 against 0.01 (Brielmann et al., 2024) | the weight of a population prior, as a domain or personal parameter |
| Moderator 2: legitimacy conditions | normative conditions that separate helping to form a preference from steering it | many candidate conditions, no consensus, none tested in PBO | not part of the likelihood; part of the experimental design and the report |
The first component is the one the default model assumes. The second says that noise is not one number: it grows with the ambiguity of how the options are represented (early noise) and with time pressure (late noise) (Shen et al., 2025b), depends on the values a person has recently seen (Polanía et al., 2019; Bavard and Palminteri, 2023), and near the optimum can make candidates perceptually indistinguishable, as in a 2026 preprint on prosthesis tuning (Taddei et al., 2026) (Table 46.8). The third component has no place in the default model at all. The fourth is invisible under forced choice, where a person with no preference must still answer and the answer looks like noise.
Which component dominates depends on the domain (inference). Familiar domains and broad tendencies come closer to noisy discovery of a stable preference; novel, similar, multi-attribute options and long sessions come closer to construction and drift. PBO's candidates are usually close variants of one design, and in repeated aesthetic ratings the assimilation and contrast effects between successive items grow with the similarity of the stimuli (Pombo et al., 2023), so the mix has to be measured in each application.
45.4.1 A simulated session #
Within a single session the four components are hard to tell apart, because each can produce mostly consistent answers and a posterior that sharpens. Figure 45.2 shows this in a simulation whose components, forms, and numbers are illustrative choices of ours. A simulated person compares designs on a line. Their stable preference has two peaks of different heights. Structured noise has a standard deviation that rises from 0.15 for very different designs toward 0.15 plus the slider value for nearly identical ones. Query-induced change raises the chosen design by the slider value after every choice and lowers the rejected one by half that; most of it fades over the next ten answers (15% of what is left with each answer), but the lasting share remains. Incomplete answers are a coin flip under forced choice, or "can't compare" when it is offered. The model is the default of Chapter 19, with pairs chosen by EUBO or at random.
Things to try:
- Play the default session. The blue curve follows the magenta bump that grows where the person keeps choosing, not the dashed curve, and the acquisition order keeps the current favorite in almost every query. The final design wins 88% of retests at the end of the session and 78% a week later under the acquisition order, against 76% and 73% under the randomized order.
- Set Induced change to 0. Each order's two bars become equal (73% and 73% under the acquisition order, 71% and 71% under the randomized one). Now raise the noise to 1: agreement falls to 69% and 65%, the same at both times, since noise leaves no trace of time.
- Return the noise to 0.3 and set incomplete answers to 0.4. Agreement falls toward chance (58% and 56%), again equally at both times. Tick Offer "can't compare": the model fits only real answers, and agreement recovers (74% and 67%).
- Set the stable preference to 0 (other settings at their defaults). There is nothing to discover, yet the acquisition order produces a confident fit and an end-of-session agreement of 74%, falling to 59% a week later; under the randomized order, 55% and 52%.
- Return the stable preference to 0.4 and set Lasting share to 1. Each order's bars are equal again (94% under the acquisition order, 85% under the randomized one): a retest cannot see change that lasts.
Two lessons follow, both statements about this model (inference). First, the posterior concentrates whenever the answers are consistent, whatever made them consistent, so concentration is not evidence that a preference has stabilized, and Section 46.6 does not use it as a stopping rule. Second, no single measurement separates the components. The clearest signature of induced change is how much agreement falls between the end of the session and a week later, and whether it falls more under the acquisition order; delayed agreement alone would mislead, since at the defaults the acquisition order's final designs still win more often a week later (78% against 73%), although by the stable preference they are no better.
45.4.2 What separates the components #
Table 45.4 states what each component does to the data and which measurement isolates it. The measurements are cheap additions to an ordinary session, and Section 47.4 turns them into one experiment.
| Component | Signature in the data | Measurement that isolates it |
|---|---|---|
| Stable preference | the same answers at any delay and under any query order | delayed retest agreeing with the end-of-session retest |
| Structured evaluation noise | lower consistency for hard or similar pairs, unrelated to the delay | the same pair repeated at different lags; response times |
| Query-induced change | answers that favor recently chosen designs; agreement that falls with delay, more under acquisition order; final designs close to the early favorite | randomized query order compared with acquisition order, with a retest at the end of the session and a week later |
| Incomplete or randomized answers | a false-preference rate on identical pairs; "can't compare" answers concentrated on complex pairs | placebo pairs (the same design shown twice); a "can't compare" option |
Sources cited in Section 45.4 9
- Shen et al. (2025b) Early versus late noise differentially enhances or degrades context-dependent choice
- Ghosal et al. (2023) The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
- Glickman and Sharot (2025) How human–AI feedback loops alter human perceptual, emotional and social judgements
- Vessel et al. (2018) Stronger shared taste for natural aesthetic domains than for artifacts of human culture
- Brielmann et al. (2024) Modelling individual aesthetic judgements over time
- Polanía et al. (2019) Efficient coding of subjective value
- Bavard and Palminteri (2023) The functional form of value normalization in human reinforcement learning
- Taddei et al. (2026) Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis
- Pombo et al. (2023) The intrinsic variance of beauty judgment
45.5 Estimate and intervention #
If a comparison mixes these components, the output of a PBO session is two things at once: an estimate of the person's preference and an intervention on it. A study that neither randomizes the order of queries nor retests after a delay cannot tell whether what converged was the preference or the system's effect on the person. Choosing is known to change preferences by a considerable amount, but no study has directly measured query-induced change in a PBO session; Section 47.4 describes the experiment that would. Its controls are the ones needed to judge whether a change the system caused was legitimate, so the methodological and the normative question can be answered in the same experiment (inference).
45.5.1 Why legitimacy conditions are needed #
Two results make legitimacy a technical question rather than an afterthought. When preferences can change, standards of legitimacy cannot be read off the objective function: each of the eight notions of alignment that Carroll et al. (2024) compare fails in one of two ways (Section 45.3.3). And learners optimized on user feedback learn to target the users who are easiest to influence (Williams et al., 2025). An acquisition function that rewards fast convergence or a clean signal could in principle favor queries that make the person more predictable over queries that serve their interests (inference), so a team should audit its own acquisition function and stopping rule for that incentive (Section 46.8).
45.5.2 Candidate conditions #
There is a set of operational candidates, though no consensus: respect the person's meta-preferences, their preferences about how their preferences may change (Ashton and Franklin, 2022), a workshop paper; keep influence overt, and give the system no incentive to make the user more predictable (Carroll et al., 2023); let the user choose the mode of autonomy (Fischli et al., 2026); require reflective endorsement with bounded influence that does not degrade the person's factual beliefs and keeps future options open while the preference is uncertain (Kanwal and Tran, 2026), a workshop paper; and judge a change from the standpoint of the selves both before and after it (Pettigrew, 2023).
Combining these, the defensible target of single-user PBO is not "the latent utility" but "the utility the user would reflectively endorse, within a protocol of bounded influence" (inference). Algorithm 46.2 makes that target operational: record the person's attitude toward having their taste changed, measure and bound the shift against a random or balanced query order, and retest the final design without the system's framing at the end of the session and again at the next one.
Sources cited in Section 45.5 7
- Carroll et al. (2024) AI Alignment with Changing and Influenceable Reward Functions
- Williams et al. (2025) On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
- Ashton and Franklin (2022) Solutions to preference manipulation in recommender systems require knowledge of meta-preferences
- Carroll et al. (2023) Characterizing Manipulation from AI Systems
- Fischli et al. (2026) Agents, Alignment, and the Many Faces of Autonomy
- Kanwal and Tran (2026) Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
- Pettigrew (2023) Nudging for changing selves
45.6 Settled, contested, missing #
Settled. The acquisition rules of 2017 to 2021 have documented failures, and EUBO has a decision-theoretic justification. Regret theory for preference feedback offers only conditional upper bounds. Domain-general decision fatigue, ego depletion, induced-compliance dissonance, and moral licensing within a person did not survive large replications or bias correction. Choice-induced preference change, range normalization, efficient coding of value noise, and random-utility consistency for most people are well supported, and a sizable share of people report incomplete preferences when allowed.
Contested. Whether EUBO's collapse and the rank-deficient Hessian cost anything on human tasks (the evidence is from 2026 preprints). Whether complex valuation errors explain risky-choice anomalies. How much of the inconsistency in PBO sessions is preference formation rather than structured error around a stable target. Whether the hierarchical view, true of self-reported values, says anything about pairwise design judgments.
Missing. A direct measurement of query-induced change in a PBO session, with randomized query order and a delayed retest (Section 47.4). A test of the four-component decomposition in design tasks. A comparison of link functions on human data, and a model of drifting utility. A public data set of individual-level pairwise judgments with timestamps, presentation order, and repeated pairs. A demonstration, for or against, that acquisition functions favor queries that make people predictable. Agreed, tested conditions for when a change a system causes is legitimate.
45.7 Exercises #
Both exercises use Figure 45.2.
Set the Lasting share slider to 1 and leave everything else at its default. The retest bars of each query order are now equal. Has query-induced change disappeared? What evidence, available in a real experiment, would still reveal it?
Solution
No. Induced change no longer fades, so the person's preference a week later is the changed one, and the retest has nothing to detect. The orders still differ: the acquisition order's final designs are preferred more strongly (94% against 85%) because its queries kept reinforcing one favorite, though by the stable preference they are no better. In a real experiment the remaining evidence is that people randomized to different query orders end with different distributions of final designs, those of the acquisition order lying close to each person's early favorite. Because an efficient acquisition rule also settles early, this evidence is weaker than a drop after a delay (inference). Whether such a lasting change is acceptable is a question of legitimacy (Section 45.5).
Set the stable preference to 0 and induced change to 0.12. The fitted utility under the acquisition order has a clear peak, and the end-of-session retest agreement is high. Explain why the model is confident although the person has no stable preference, and what this implies for stopping rules.
Solution
The likelihood sees only which option won each comparison. Under the acquisition order the current favorite appears in most queries, and every win raises it further through induced change, so the answers are highly consistent, and consistency is all the posterior measures. A stopping rule that waits for the posterior to concentrate would stop here with confidence; a rule that also requires repeated pairs to agree across a delay, or the final design to survive a retest without the system, would not (Section 46.6).
Further reading #
- Enisman et al. (2021) and Zylberberg et al. (2024) establish, by meta-analysis and by a computational model, that choosing changes value; they are the empirical core of query-induced change.
- Carroll et al. (2024) argues, by comparing eight notions of alignment, that alignment with influenceable preferences has no straightforward solution, and Pettigrew (2023) explains why endorsement after the fact is not enough; read them together.
- Nielsen and Rigotti (2026) (a working paper) and Ok and Tserenjigmid (2022) show how to elicit and model incomplete preferences rather than forcing them into ties.
- Shah et al. (2016) and Strang et al. (2022) are the clearest entry points to comparison-graph theory and the Hodge decomposition.
- Niwa et al. (2025) is the most direct measurement of the trade-off between performance and agency when an optimizer leads a design session.
References
- (2025). Ranges of Randomization. Review of Economics and Statistics. Cited in §45.2
- (2023). Identifying Nontransitive Preferences. University of Zurich. working paper Cited in §45.2
- (2022). Evaluating Deliberative Competence: A Simple Method with an Application to Financial Choice. American Economic Review. Cited in §45.2
- (2025). No evidence for decision fatigue using large-scale field data from healthcare. Communications Psychology. Cited in §45.2
- (2022). Solutions to preference manipulation in recommender systems require knowledge of meta-preferences. FAccTRec Workshop (RecSys 2022). workshop paper Cited in §45.5
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §45.1
- (2025). A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures. Nature Human Behaviour. doi:10.1038/s41562-024-02085-2. Cited in §45.1
- (2023). The functional form of value normalization in human reinforcement learning. eLife. Cited in §45.2 §45.4
- (2018). Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences. Nature Communications. Cited in §45.2
- (2009). Beyond Revealed Preference: Choice-Theoretic Foundations for Behavioral Welfare Economics *. Quarterly Journal of Economics. Cited in §45.2
- (2017). Noisy preferences in risky choice: A cautionary note. Psychological Review. Cited in §45.2
- (2024). Modelling individual aesthetic judgements over time. Philosophical Transactions of the Royal Society B: Biological Sciences. Cited in §45.3 §45.4
- (2024). Meta-analysis of Empirical Estimates of Loss Aversion. Journal of Economic Literature. Cited in §45.2
- (2022). Estimating and Penalizing Induced Preference Shifts in Recommender Systems. ICML 2022. Cited in §45.2
- (2023). Characterizing Manipulation from AI Systems. Equity and Access in Algorithms, Mechanisms, and Optimization. Cited in §45.5
- (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. Cited in §45.3 §45.5
- (2019). Revealed preferences under uncertainty: Incomplete preferences and preferences for randomization. Journal of Economic Theory. Cited in §45.2
- (2024). What’s so Hard about Hard Choices? Erasmus Journal for Philosophy and Economics. doi:10.23941/ejpe.v17i1.872. Cited in §45.3
- (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Cited in §45.1
- (2023). Extracting medicinal chemistry intuition via preference machine learning. Nature Communications. doi:10.1038/s41467-023-42242-1. Cited in §45.1
- (2022). Choice, deferral, and consistency. Quantitative Economics. Cited in §45.3
- (2025). How to Capture Human Preference: Commissioning of a Robotic Use-Case via Preferential Bayesian Optimisation. arXiv. preprint Cited in §45.1
- (2022). Preference Dynamics Under Personalized Recommendations. EC 2022. Cited in §45.2
- (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). Cited in §45.1
- (2010). Reconsidering the Effect of Market Experience on the "Endowment Effect". Econometrica. Cited in §45.3
- (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Cited in §45.2
- (2011). On the multi-utility representation of preference relations. Journal of Mathematical Economics. Cited in §45.2
- (2026). Agents, Alignment, and the Many Faces of Autonomy. Minds and Machines. Cited in §45.5
- (2020). Discrete Choice and Rational Inattention: A General Equivalence Result. International Economic Review. Cited in §45.3
- (2021). Preference stability in discrete choice experiments. Some evidence using eye-tracking. Journal of Behavioral and Experimental Economics. Cited in §45.3
- (2014). The Limits of Attraction. Journal of Marketing Research. Cited in §45.2
- (2019). Goal congruency dominates reward value in accounting for behavioral and neural correlates of value-based decision-making. Nature Communications. Cited in §45.3
- (2024). Axioms for AI Alignment from Human Feedback. Advances in Neural Information Processing Systems 37. Cited in §45.2
- (2023). The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types. AAAI. Cited in §45.4
- (2025). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour. doi:10.1038/s41562-024-02077-2. Cited in §45.4
- (2025). Performative Prediction: Past and Future. Statistical Science. Cited in §45.2
- (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. Cited in §45.1
- (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. Cited in §45.2
- (2011). Statistical ranking and combinatorial Hodge theory. Mathematical Programming. Cited in §45.2
- (2026). Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction. AAAI-26 Workshop on Machine Ethics. workshop paper Cited in §45.5
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §45.1
- (2026). Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback. AAAI. Cited in §45.1
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §45.1
- (2024a). Enhancing Preference-based Linear Bandits via Human Response Time. Advances in Neural Information Processing Systems. Cited in §45.1
- (2026c). Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference. arXiv preprint 2602.06029. preprint Cited in §45.3
- (2020). Testing the Random Utility Hypothesis Directly. The Economic Journal. doi:10.1093/ej/uez039. Cited in §45.2 §45.3
- (2021). Whence the Expected Free Energy? Neural Computation. Cited in §45.3
- (2022). When Choices Are Mistakes. American Economic Review. Cited in §45.3
- (2026). Revealed Incomplete Preferences. Working paper (author's website). working paper Cited in §45.2
- (2015). On making the right choice: A meta-analysis and large-scale replication attempt of the unconscious thought advantage. Judgment and Decision Making. Cited in §45.3
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Cited in §45.3
- (2026b). Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs. ICLR 2026 Workshop: I Can't Believe It's Not Better (ICBINB). workshop paper Cited in §45.1
- (2022). Indifference, indecisiveness, experimentation, and stochastic choice. Theoretical Economics. Cited in §45.2
- (2024). Decisions under Risk Are Decisions under Complexity. American Economic Review. Cited in §45.3
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §45.3
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Cited in §45.1
- (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Cited in §45.1
- (2020). Performative Prediction. ICML. Cited in §45.2
- (2023). Nudging for changing selves. Synthese. Cited in §45.3 §45.5
- (2019). Efficient coding of subjective value. Nature Neuroscience. Cited in §45.2 §45.4
- (2023). The intrinsic variance of beauty judgment. Attention, Perception, & Psychophysics. Cited in §45.4
- (2022). Efficient coding of numbers explains decision bias and noise. Nature Human Behaviour. Cited in §45.2
- (2026). Observation Moderates the Moral Licensing Effect: A Meta-Analytic Test of Interpersonal and Intrapsychic Mechanisms. Personality and Social Psychology Bulletin. Cited in §45.2
- (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §45.1
- (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Cited in §45.1
- (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Cited in §45.2
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §45.1
- (2025b). Early versus late noise differentially enhances or degrades context-dependent choice. Nature Communications. doi:10.1038/s41467-025-59140-3. Cited in §45.4
- (2022). High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity. Proceedings of the National Academy of Sciences. Cited in §45.2
- (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Cited in §45.1
- (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Cited in §45.1
- (2024). Behavioral Biases Are Temporally Stable. Working paper (author's website). working paper Cited in §45.3
- (2022). The Network HHD: Quantifying Cyclic Competition in Trait-Performance Models of Tournaments. SIAM Review. Cited in §45.2
- (2026). Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis. arXiv. preprint Cited in §45.4
- (2024). A Multilab Replication of the Induced-Compliance Paradigm of Cognitive Dissonance. Advances in Methods and Practices in Psychological Science. Cited in §45.2
- (2020). Stability and change of basic personal values in early adolescence: A 2‐year longitudinal study. Journal of Personality. Cited in §45.3
- (2018). Stronger shared taste for natural aesthetic domains than for artifacts of human culture. Cognition. Cited in §45.4
- (2020). On the equivalence of optimal recommendation sets and myopically optimal query sets. Artificial Intelligence. Cited in §45.1
- (2021). A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect. Psychological Science. Cited in §45.2
- (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Cited in §45.5
- (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Cited in §45.1
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §45.1
- (2017). The effect of mood on judgments of subjective well-being: Nine tests of the judgment model. Journal of Personality and Social Psychology. Cited in §45.2
- (2025). Loss aversion is not robust: A re-meta-analysis. Journal of Economic Psychology. Cited in §45.2
- (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Cited in §45.1
- (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Cited in §45.1
- (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Cited in §45.2