Bayesian Optimization
Part IX: What Is a Preference?
中文

Neuroscience and Computational Cognitive Science

The two previous chapters reported behavior: what people choose, how consistently, and under which conditions (Chapter 37, Chapter 38). This chapter asks what the brain and computational models of the mind add for someone who builds or evaluates a preferential Bayesian optimization (PBO) system: whether a single scale of value exists inside the head, as the latent utility of PBO presumes; what a response time reveals beyond the answer; and where the noise in a comparison comes from, and whether it stays fixed.

The main change between 2017 and 2026 was to treat value as constructed during a decision and revised by the choice itself. A scalar value has causal support within one task, but a cardinal "common currency" across tasks remains disputed. The noise in value judgments depends on the values seen before, so discrimination does not decline near the optimum for value judgments, only for perceptual parameters. Attention moves choice causally, but a little. Response times entered preference learning with guarantees, though not yet in live experiments with people. Neural signals add only about 4 percentage points over self-report for an individual, and active inference is not yet a usable alternative to PBO. The centerpiece is the drift-diffusion model of Section 39.3, which you can run in Figure 39.1.

39.1 Neuroeconomics and value-based decisions #

Neuroeconomics looks for the brain's representation of value. A common claim holds that the ventromedial prefrontal cortex (vmPFC) and orbitofrontal cortex (OFC), regions behind the forehead and above the eyes, integrate the values of different kinds of reward into a common currency, a single scale on which food, money, and music can be compared (Levy and Glimcher, 2012; Bartra et al., 2013). A critique, published in Behavioral Neuroscience, argues that value-related signals may reflect salience, arousal, or attention, that OFC coding adapts to the range of values on offer, and that choice can bypass value altogether through direct learning of action policies; its central claim is that a cardinal common-currency value is not, by default, represented in the brain and used for choice (Hayden and Niv, 2021).

Within a task, the causal evidence grew stronger. Electrically stimulating the OFC of macaque monkeys shifted their choices by raising the value of an individual offer (Ballesta et al., 2020), and in intracranial recordings from 36 patients with epilepsy, subjective value could be decoded from vmPFC and lateral OFC activity, with both a linear signal (value) and a quadratic one (confidence) (Lopez-Persem et al., 2020). OFC neurons adapt to the range of values on offer, but only partially, and in a linear decision model the resulting raised floor of activity increases choice variability (Conen and Padoa-Schioppa, 2019). Value is also constructed during the decision. A metacognitive control model, in which a fast, uncertain initial estimate of value is refined under a trade-off between mental effort and confidence, predicts at once the response time, confidence, changes of mind, and choice-induced preference change, and its authors argue that a preference reversal can be a re-evaluation rather than noise (Lee and Daunizeau, 2021). The size of choice-induced preference change is in Section 37.2.2.

What this means for PBO. The evidence supports a scalar latent utility within one task and one session, identifiable only up to a monotone transformation (Section 18.4), in line with the argument that preference learning needs only order consistency (Sun et al., 2025); it does not support reusing a learned scale in another task without recalibration (inference). Because coding partly adapts to the range presented, overly wide exploratory queries early on may leave higher variability later (inference). Confidence is encoded together with value, so scaling the noise of the probit link by the confidence of each comparison is a direct way to implement heteroscedastic noise (inference). All of this evidence comes from food, consumer goods, juice offers, and money, so for continuous design parameters these are working hypotheses (Section 39.9).

Sources cited in Section 39.1 8
  1. Levy and Glimcher (2012) The root of all value: a neural common currency for choice
  2. Bartra et al. (2013) The valuation system: A coordinate-based meta-analysis of BOLD fMRI experiments examining neural correlates of subjective value
  3. Hayden and Niv (2021) The case against economic values in the orbitofrontal cortex (or anywhere else in the brain)
  4. Ballesta et al. (2020) Values encoded in orbitofrontal cortex are causally related to economic choices
  5. Lopez-Persem et al. (2020) Four core properties of the human brain valuation system demonstrated in intracranial signals
  6. Conen and Padoa-Schioppa (2019) Partial Adaptation to the Value Range in the Macaque Orbitofrontal Cortex
  7. Lee and Daunizeau (2021) Trading mental effort for confidence in the metacognitive control of value-based decision-making
  8. Sun et al. (2025) Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

39.2 Neuromarketing and neuroforecasting #

A common claim holds that neural signals reveal preferences people cannot report. Its best-known support is a study in which activity in the nucleus accumbens (part of the ventral striatum) of adolescents listening to songs predicted the songs' later sales while their own ratings did not (Berns and Moore, 2012). The evidence base is small: after exclusions there were 27 adolescents; 87 of the 120 songs had sales data; accumbens activity correlated with log sales at R=0.32R = 0.32 (p=0.004p = 0.004), the mean likability rating at R=0.110R = 0.110 (not significant); a logistic regression classified only 30% of the hits correctly; and the data came from an experiment originally designed to study social influence (Berns and Moore, 2012). We found no independent direct replication. A later paper reporting 97% accuracy in classifying hit songs (Merritt et al., 2023) was audited: the data had been oversampled before they were split into training and test sets, covered only 24 songs and 33 listeners, and the corrected accuracy was hardly better than chance (Kapoor and Narayanan, 2023), a non-peer-reviewed audit.

The sturdier progress is at the level of populations. In 30 participants scanned with functional magnetic resonance imaging (fMRI), nucleus accumbens activity predicted the outcomes of internet crowdfunding campaigns weeks later while the participants' own choices did not, and the result replicated in a second study (Genevsky et al., 2017). With internet samples of 2,956 and 992 people, forecasts based on behavior varied with the demographic representativeness of the laboratory sample, while forecasts based on brain activity stayed significant (Genevsky et al., 2025). For individuals the gain is limited: combining several electroencephalography (EEG) measures with machine learning predicted each person's most and least liked of six food products with 68.5% accuracy, 4.07 percentage points above self-report alone (Hakim et al., 2021). We found no meta-analysis of the accuracy of neural prediction, and the out-of-sample evidence comes mainly from one laboratory.

What this means for PBO. PBO learns one person's utility, while the advantage of neural forecasting is documented for aggregates, so its most defensible use is a population prior or a warm start, not a personal likelihood (inference). A session of a few dozen comparisons cannot train a personal decoder, and any PBO system that uses a neural proxy needs strict held-out evaluation (inference).

Sources cited in Section 39.2 6
  1. Berns and Moore (2012) A neural predictor of cultural popularity
  2. Merritt et al. (2023) Accurately predicting hit songs using neurophysiology and machine learning
  3. Kapoor and Narayanan (2023) No, you still cannot predict hit songs using machine learning
  4. Genevsky et al. (2017) When Brain Beats Behavior: Neuroforecasting Crowdfunding Outcomes
  5. Genevsky et al. (2025) Neuroforecasting reveals generalizable components of choice
  6. Hakim et al. (2021) Machines learn neuromarketing: Improving preference prediction from self-reports using multiple EEG measures and machine learning

39.3 Sequential sampling and response times #

A common claim, drawing on decision field theory (Busemeyer and Townsend, 1993) and the attentional drift-diffusion model (Krajbich et al., 2010), treats response time as a natural measure of preference uncertainty: a fast answer signals a large difference in value, a slow one a small difference. Two theoretical results back it: in preference-based bandits, response times are most informative for easy queries (Li et al., 2024a), and choices and response times together can identify the distribution of latent preferences at several points, according to a 2026 working paper (Benkert et al., 2026).

39.3.1 The drift-diffusion model #

The drift-diffusion model (DDM) describes a two-option decision as noisy evidence accumulating over time (Ratcliff, 1978). Think of a counter that starts at zero. Every few milliseconds it moves up a little if the sample of evidence favors option A and down if it favors B; on average it drifts at rate vv, the drift rate, but each step is noisy. The decision is made when the counter first reaches +a+a (choose A) or −a-a (choose B); the boundary aa expresses caution. The response time is the time to reach a boundary plus a fixed non-decision time t0t_0 for perceiving and pressing a key. In value-based choice the drift is proportional to the difference in value, v=k Δuv = k\,\Delta u with Δu=u(A)−u(B)\Delta u = u(A) - u(B).

Two closed forms describe the model with unit noise (Bogacz et al., 2006). The probability of choosing A and the mean decision time are

P(choose A)=11+e−2av,E[T]=avtanh⁡(av).\Prob(\text{choose A}) = \frac{1}{1 + e^{-2 a v}}, \qquad \E[T] = \frac{a}{v} \tanh(a v).
(39.1)

The first is the logistic function of 2akΔu2 a k \Delta u. A model of the decision process, in other words, produces the logistic (Bradley-Terry) link of Section 16.4 as its choice probability, with a slope set jointly by the person's caution aa and the drift scale kk.

Derivation Why the drift-diffusion model gives the logistic link
  1. Let h(x)h(x) be the probability of reaching +a+a before −a-a when the counter is at xx. Over a short time dtdt the counter moves by v dtv\,dt plus noise of variance dtdt, so h(x)=E[h(x+v dt+dt Z)]h(x) = \E[h(x + v\,dt + \sqrt{dt}\,Z)] with Z∼N(0,1)Z \sim \N(0, 1).
  2. Expand hh to second order and take the expectation, using E[Z]=0\E[Z] = 0 and E[Z2]=1\E[Z^2] = 1: h(x)=h(x)+v h′(x) dt+12h′′(x) dth(x) = h(x) + v\,h'(x)\,dt + \tfrac12 h''(x)\,dt, so 12h′′+v h′=0\tfrac12 h'' + v\,h' = 0.
  3. The general solution of this linear equation is h(x)=A+B e−2vxh(x) = A + B\,e^{-2 v x}, as substituting it confirms.
  4. The boundaries fix the constants: h(−a)=0h(-a) = 0 and h(a)=1h(a) = 1 give h(x)=1−e−2v(x+a)1−e−4vah(x) = \dfrac{1 - e^{-2v(x + a)}}{1 - e^{-4 v a}}.
  5. At the start, x=0x = 0. Factor the denominator as 1−e−4va=(1−e−2va)(1+e−2va)1 - e^{-4va} = (1 - e^{-2va})(1 + e^{-2va}) and cancel: h(0)=11+e−2vah(0) = \dfrac{1}{1 + e^{-2va}}, the logistic function of 2av2av.

The model makes the claim about response times precise. When Δu=0\Delta u = 0 the drift vanishes, choices are coin flips, and the mean decision time takes its largest value, a2a^2. As ∣Δu∣|\Delta u| grows, the choice probability saturates near 0 or 1 while the response time keeps falling. Beyond a moderate utility difference the choice says almost nothing new, but the response time still does, which is why response times help most for easy queries (Li et al., 2024a).

−2−1012evidence for Achoose Achoose B0.00.51.01.52.02.53.0decision time (s)0.00.20.40.60.81.0P(choose A)−101utility difference Δu0.00.51.01.5mean response time (s)−101utility difference ΔuP(choose A) = 0.77 · mean response time 1.20 s
−2−1012evidence for Achoose Achoose B0123decision time (s)0.00.20.40.60.81.0P(choose A)−101utility difference Δu0.00.51.01.5mean response time (s)−101utility difference ΔuP(choose A) = 0.77 · mean response time 1.20 s
Figure 39.1 The drift-diffusion model. Top: sample paths of accumulated evidence for option A over option B, ending at the upper boundary (choose A, blue) or the lower one (choose B, orange); the strips show the decision times of 600 simulated decisions for each answer, leaving out decisions still undecided after 3 seconds. Bottom: the closed-form probability of choosing A, a logistic function of the utility difference, and the mean response time, which peaks at Δu = 0; the dot marks the current pair. The overall-value effect multiplies the drift by 1 + 1.5V, an illustrative stand-in for the finding that pairs of good options are decided faster and more accurately; the dashed curves show the same pair at V = 0. Drift scale, non-decision time (0.3 s), and the size of the value effect are illustrative.

Things to try:

  • Set Δu=0\Delta u = 0. About half of the simulated decisions end at each boundary, and the decision-time distributions are widest.
  • Raise Δu\Delta u to 1 and beyond. The choice probability is already near 1, while the mean response time keeps dropping: in this range only the response time distinguishes a large utility difference from a very large one.
  • Raise the caution aa. Choices become more accurate and slower, and the logistic curve steepens: the slope of the link, which a comparison model absorbs into its noise scale σ\sigma, reflects the person's caution as well as their utility.
  • Switch on the overall-value effect and raise VV. The same utility difference now yields a faster and more accurate choice, and a model that ignores overall value would read the response as a larger utility gap.

39.3.2 Attention and response times, measured #

The causal effect of attention now has a number, which Figure 39.2 shows. Pooling experiments that manipulated attention in two-option preferential choice, increasing the total time an option was looked at raised the probability of choosing it to P=.541P = .541 (95% confidence interval [.523,.560][.523, .560]), controlling which option was looked at last gave P=.532P = .532 ([.518,.547][.518, .547]), and manipulating the first fixation had no effect (P=.507P = .507, [.497,.516][.497, .516], p=.18p = .18); slight publication bias lowers these values a little (Bhatnagar and Orquin, 2022). The discount on the unattended option is not a constant. In the study that introduced the attentional model, 0.3 was the best fit to the pooled data of 39 participants, while fits to each participant gave a mean of 0.52 with a standard deviation of 0.3 (Krajbich et al., 2010); in a preregistered study (N=61N = 61), gaze effects weakened when the overall value of the options was high (Ting and Gluth, 2025).

Probability of choosing the option attention was steered toward.480.500.520.540.560.580probability (dashed line at 0.5: no effect)Total looking time.541 [.523, .560]Last fixation.532 [.518, .547]First fixation.507 [.497, .516], p = .18In 40 comparisons that all carried this manipulation, the option attention was steeredtoward would win about 1.6 more answers (0.9 to 2.4) than with no effect.Discount on the option not being looked at0.00.20.40.60.81.0discount (1: the unattended option counts fully)pooled fit, 39 participantseach participant, mean ± 1 SD0.30.52 ± 0.3
Probability of choosing the option attentionwas steered toward.480.500.520.540.560.580probability (dashed: no effect)Total looking timeLast fixationFirst fixation (p = .18)In 40 comparisons that all carried thismanipulation, the option attention wassteered toward would win about 1.6 moreanswers (0.9 to 2.4) than with no effect.Discount on the option not being looked at0.00.20.40.60.81.0discount (1: counts fully)pooled fit, 39 participantseach participant, mean ± 1 SD0.30.52 ± 0.3
Figure 39.2 How much attention moves choice. Top: the pooled probability of choosing an option when an experiment steered attention toward it, with 95% confidence intervals, from a meta-analysis of two-option preferential choice (Bhatnagar and Orquin, 2022); the dashed line at 0.5 is no effect, and slight publication bias would lower each value a little. The readout converts the selected row into answers out of a session in which every comparison carried that manipulation, which is arithmetic on the pooled values, not a further finding. Bottom: the discount on the option not being looked at in the attentional drift-diffusion model, fitted to the pooled data of 39 participants and to each participant separately (mean and one standard deviation) (Krajbich et al., 2010).

Things to try:

  • Compare the three intervals with the 0.5 line. Those for total looking time (.523 to .560) and for the last fixation (.518 to .547) lie above it; the one for the first fixation (.497 to .516) contains it.
  • Select total looking time and set the session to 40 comparisons. If every comparison carried that manipulation, the option looked at longer would win about 1.6 more answers (0.9 to 2.4) out of 40 than with no effect.
  • In the lower panel, the pooled discount, 0.3, lies below the mean of the individual fits, 0.52, whose standard deviation, 0.3, is as large as the pooled value itself.

Response time measures preference strength, with confounds. In three domains of choice, decisions between two high-value options were mostly faster and more accurate, which in the DDM means a higher drift rate, contrary to diminishing sensitivity to value (Shevlin et al., 2022); the negative relation between overall value and response time replicated in Ting and Gluth (2025). Response-time proxies for preference strength are often noisy and confounded (Kaufmann et al., 2025).

39.3.3 Response times enter preference learning #

The main development of 2017 to 2026 is that response times entered preference learning with guarantees, Gaussian processes included. Shvartsman et al. (2024) approximated the DDM likelihood differentiably with a family of skewed three-parameter distributions, so that response times can enter a Gaussian process model of binary choice, and report better estimates of latent value and predictions of held-out choices on three real psychophysics and preference data sets (UAI 2024; Section 27.2 gives the details). Sawarni et al. (2025) built a loss that uses response times under the EZ-diffusion model, a simplified DDM whose parameters can be computed in closed form from the mean and variance of response times and the accuracy (Wagenmakers et al., 2007); for linear rewards it turns the error of standard preference learning, which grows exponentially with the magnitude of the reward, into polynomial growth, and it extends to nonparametric reward spaces. A July 2026 preprint analyzes each pairwise label under rational inattention, the idea that people pay for information and attend only as much as it is worth (Section 40.5). It argues that heterogeneous attention can lead a Bradley-Terry model to recover a misleading ranking, and that in perceptual comparisons response times and gaze carry information about the size of the gap that labels lack (Xing, 2026).

What this means for PBO. A joint likelihood of choice and response time can already be used with a Gaussian process surrogate. It should take overall value (or a proxy, such as the current posterior means of the two options) as a covariate of the drift rate or the boundary; otherwise fast answers late in a session, when both options are good, will be read as large utility differences (inference), the misreading Figure 39.1 shows. Simultaneous, side-by-side presentation should be preferred, and options that must be experienced in sequence (audio, touch, gait) should have their order counterbalanced, with an order bias term in the likelihood; the expected presentation effect is a few percentage points of choice probability (inference). The spread of individual discounts supports per-user bias or discount parameters rather than a population constant such as 0.3 (inference). The existing response-time Gaussian process work used offline data; the experiments that would test these recommendations with people in the loop are listed in Section 39.9.

Sources cited in Section 39.3 14
  1. Busemeyer and Townsend (1993) Decision field theory: A dynamic-cognitive approach to decision making in an uncertain environment
  2. Krajbich et al. (2010) Visual fixations and the computation and comparison of value in simple choice
  3. Li et al. (2024a) Enhancing Preference-based Linear Bandits via Human Response Time
  4. Benkert et al. (2026) Time is Knowledge: What Response Times Reveal
  5. Ratcliff (1978) A theory of memory retrieval
  6. Bogacz et al. (2006) The physics of optimal decision making: A formal analysis of models of performance in two-alternative forced-choice tasks
  7. Bhatnagar and Orquin (2022) A meta-analysis on the effect of visual attention on choice
  8. Ting and Gluth (2025) High overall values mitigate gaze-related effects in perceptual and preferential choices
  9. Shevlin et al. (2022) High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity
  10. Kaufmann et al. (2025) ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning
  11. Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences
  12. Sawarni et al. (2025) Preference Learning with Response Time: Robust Losses and Guarantees
  13. Wagenmakers et al. (2007) An EZ-diffusion model for response time and accuracy
  14. Xing (2026) Attention Limited Reward Learning

39.4 Efficient coding and normalization #

The Weber argument holds that as the optimizer converges on the high-utility region, the person's ability to tell options apart declines, so queries should target regions where the expected utility difference exceeds the just-noticeable difference (the smallest difference a person reliably detects). A related argument treats comparison noise as fixed expression noise around a stable utility, with rational inattention (Sims, 2003; Matějka and McKay, 2015) as its microfoundation.

For value judgments, the Weber argument fails: decisions between high-value options are faster and more accurate (Section 39.3.2). Efficient coding, the principle that a system with limited capacity should spend its resolution where inputs are most frequent, predicts that discrimination tracks the density of values in the environment rather than their magnitude. Polanía et al. (2019) had 38 participants rate and choose among 64 familiar foods, the first of four experiments with 127 participants in all: a valuation model that maximizes information under limited coding resources explained rating variability, choice consistency, and confidence together, with choices predicted to be more consistent at values where the prior density is high. Related work found that rarer numbers are coded more noisily (Prat-Carrabin and Woodford, 2022), that in two preregistered experiments risk-taking was more sensitive to payoffs that occurred more often (Frydman and Jin, 2022), and that monkeys chose more sensitively in environments with lower reward variance (Zimmermann et al., 2018). The noise itself changes with context and history, then, rather than being fixed expression noise around a stable utility.

Divisive normalization, in which each option's value is divided by a measure of the total value on offer, remains contested. A discrete choice model with divisive normalization fitted effects of choice-set size and composition better than alternatives including range normalization (Webb et al., 2021), but a distractor effect attributed to normalization failed to replicate (Section 37.2.3), and in tasks with learned values range normalization won (Section 37.3.1). A distinction between early and late noise offers a reconciliation: when the representation of the options is blurry (early noise), a third option improves discrimination, and when time pressure dominates (late noise), the context harms it; a divisive normalization model with both kinds of noise fitted best (Shen et al., 2025b).

What this means for PBO. An acquisition function that concentrates queries near the incumbent narrows the range of utilities the person experiences. Efficient coding predicts that sensitivity in that region rises during the session, partially and with a lag, and a probit likelihood with a fixed noise scale would misread that late precision as a steeper utility surface (inference). A noise scale that depends on how typical the options' utilities are, relative to those presented so far, can represent this; a simpler practice is to insert a few fixed reference pairs from rarely visited regions and measure discrimination there directly (inference). Whether Weber-style compression or efficient coding dominates in a domain can be checked with a few repeated pairs at different stages of a session: for value judgments the evidence favors efficient coding, while perceptual parameters (stiffness, torque, color) have floors (Section 39.7) (inference).

Sources cited in Section 39.4 8
  1. Sims (2003) Implications of rational inattention
  2. Matějka and McKay (2015) Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model
  3. Polanía et al. (2019) Efficient coding of subjective value
  4. Prat-Carrabin and Woodford (2022) Efficient coding of numbers explains decision bias and noise
  5. Frydman and Jin (2022) Efficient Coding and Risky Choice
  6. Zimmermann et al. (2018) Multiple timescales of normalized value coding underlie adaptive choice behavior
  7. Webb et al. (2021) The Normalization of Consumer Valuations: Context-Dependent Preferences from Neurobiological Constraints
  8. Shen et al. (2025b) Early versus late noise differentially enhances or degrades context-dependent choice

39.5 Predictive processing and active inference #

The free-energy principle proposes that brains, and perhaps all self-organizing systems, act to minimize variational free energy, a bound on how surprising their sensations are under their internal model (Friston, 2010). Active inference applies it to action: an agent chooses actions that minimize expected free energy, and preferences are prior beliefs about the outcomes the agent expects to observe. Expected free energy decomposes into a pragmatic value (reaching preferred outcomes) and an epistemic value (gaining information) (Friston et al., 2015). The decomposition has been called "exactly" the balance of exploration and exploitation in PBO, with an exploration term "derived from first principles rather than added by hand" like the weight of an upper confidence bound (Section 12.4).

That claim needs qualification. A functional argued to be the natural extension of variational free energy into the future actively suppresses exploration, so exploration does not follow directly from minimizing free energy (Millidge et al., 2021). The first theoretical guarantee for agents that minimize expected free energy, a 2026 preprint, shows that "sufficient curiosity" ensures both posterior consistency and bounded cumulative regret, while too little curiosity leads to myopic exploitation and too much to unnecessary exploration and regret (Li et al., 2026c). The weight, in other words, still has to be set, just like the exploration weight of an upper confidence bound. The foundations are contested too: Markov blankets, the boundaries that separate a system from its environment in the theory, are persistently conflated between a tool of inference and physical boundaries (Bruineberg et al., 2021), and the "dark room problem", that an agent minimizing surprise should seek a dark, quiet room, is a real difficulty for the idea that preference is prediction (Klein, 2018). Three 2026 preprints combine active inference with Bayesian optimization: the regret guarantee above; "pragmatic curiosity", which scores queries by the information gained about task-relevant latent variables plus a potential of expected regret (Li et al., 2026e); and BOBA, an acquisition function that adds lookahead uncertainty about how a time-varying objective will change, which improved regret on synthetic dynamic benchmarks (Kelly et al., 2026).

What this means for PBO. As of September 2026, active inference is not a usable alternative to PBO (inference, from the studies above). Every concrete implementation in Bayesian optimization reduces to an acquisition function, a pragmatic term plus a weighted information-gain term, close to existing information-theoretic acquisition functions, and it replaces neither the Gaussian process preference model nor the pairwise likelihood. The abstracts of the three preprints report no experiments with human pairwise feedback (we did not read the full texts), and no study compares an expected-free-energy acquisition function with EUBO-type acquisition functions (Section 19.4) on human preference data. And the two frameworks mean different things by preference: in active inference it is the agent's prior over its own observations, in PBO the person's utility is a hidden state to be inferred, so rewriting PBO as active inference re-encodes the objective without adding any constraint on the structure of human preference (inference). One role could be useful: modeling the person as an agent whose prior preferences are learned and updated (Sajid et al., 2021) would give a generative model of preference drift, though we found no human preference data fitted that way (inference).

Sources cited in Section 39.5 9
  1. Friston (2010) The free-energy principle: a unified brain theory?
  2. Friston et al. (2015) Active inference and epistemic value
  3. Millidge et al. (2021) Whence the Expected Free Energy?
  4. Li et al. (2026c) Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference
  5. Bruineberg et al. (2021) The Emperor's New Markov Blankets
  6. Klein (2018) What do predictive coders want?
  7. Li et al. (2026e) Pragmatic Curiosity: A Unified Framework for Hybrid Learning and Optimization via Active Inference
  8. Kelly et al. (2026) BOBA: Dynamic Bayesian Optimization through Bayesian Active Inference
  9. Sajid et al. (2021) Active Inference: Demystified and Compared

39.6 Psychopharmacology and addiction #

Pharmacology offers what look like the strongest counterexamples to stable preference. A common claim distinguishes "wanting", attributed to mesolimbic dopamine, from "liking", attributed to opioid hedonic hotspots in the brain (Berridge and Robinson, 1998), and proposes modeling the two with a multi-output Gaussian process. Evidence that the two separate in human self-report is weak. In a double-blind within-subject study (n=27n = 27), levodopa raised and risperidone lowered both musical pleasure and music-related motivation together (Ferreri et al., 2019); in 131 volunteers, the dopamine antagonist amisulpride and the opioid antagonist naltrexone both reduced physical effort, but neither changed subjective ratings of wanting or liking (Korb et al., 2020). Dopaminergic drugs do change behavior. In a cross-sectional study of 3,090 treated patients with Parkinson's disease, impulse control disorders were more common among patients taking a dopamine agonist than among those not taking one (17.1% against 6.9%, odds ratio 2.72) (Weintraub et al., 2010), and in 31 healthy men levodopa weakened directed exploration (Chakroun et al., 2020).

What this means for PBO. Pairwise choices and ratings may not separate wanting from liking, so a two-output model would need different observation channels, such as effort for wanting and ratings during consumption for liking (inference). Medication state is a slow context variable: studies of PBO in health and rehabilitation should record it and allow for drift over weeks that need not be monotone, and since exploration and noise vary with dopamine state, noise parameters should be fixed per session rather than per person (inference). When a drug changes what a person prefers, which preference counts is a question of authority; in clinical settings a reflective endorsement check, made away from the peak effect of a drug or at follow-up, is a reasonable safeguard (inference), one Section 41.2 discusses. No study has elicited pairwise preferences under a pharmacological manipulation.

Sources cited in Section 39.6 5
  1. Berridge and Robinson (1998) What is the role of dopamine in reward: hedonic impact, reward learning, or incentive salience?
  2. Ferreri et al. (2019) Dopamine modulates the reward experiences elicited by music
  3. Korb et al. (2020) Dopaminergic and opioidergic regulation during anticipation and consumption of social and nonsocial rewards
  4. Weintraub et al. (2010) Impulse Control Disorders in Parkinson Disease: A Cross-Sectional Study of 3090 Patients
  5. Chakroun et al. (2020) Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making

39.7 Motor learning and adaptation #

A common claim holds that PBO works well where preferences are experiential and users have real expertise. The claim about expertise needs refining: in an ankle exoskeleton, users the study calls knowledgeable (the two groups differed in technical background) preferred higher torque than naive users, and naive users' preferred torque rose over the course of the experiment, so knowledge and exposure change the preferred setting itself, not just the noise around it (Ingraham et al., 2022).

People adapt on several time scales while a device is tuned for them. In Poggensee and Collins (2021), after training with moderate variability, personalized assistance reduced metabolic cost by 39% relative to the exoskeleton switched off, training contributed about half of the benefit and personalization about a quarter, and becoming an expert user took about 109 minutes of assisted walking. When people met a new exoskeleton context, the variability of step frequency, ankle angle range, and muscle activity first rose and then fell, on time scales that differed between variables (Abram et al., 2022). Preferences are repeatable but change: when 12 naive and 12 knowledgeable participants tuned the torque magnitude and timing of an ankle exoskeleton themselves, their preferences ranged from 7.9 to 19.4 N·m and from 54.1% to 59.2% of the gait cycle, with trial-to-trial standard deviations of 1.7 N·m and 1.5%, and convergence within a trial took 105 seconds (Ingraham et al., 2022). The preferred stiffness of a prosthetic ankle maximized the kinematic symmetry between the prosthetic and the intact joint and was not significantly related to body weight or metabolic rate (Clites et al., 2021). Perceptual resolution differs by a factor of about five between judging the stiffness of a prosthetic ankle and that of an ankle exoskeleton (inference, from two studies with different devices and people): eight people with below-knee amputations detected a 7.7% change in stiffness with 75% accuracy (Shepherd et al., 2018), while walking in an ankle exoskeleton the just-noticeable difference was about 42% for stiffness (Maberry and Martin, 2026). The optimizer and the person also adapt to each other: in three online experiments on co-adaptation, when the machine used a policy gradient, human behavior moved toward the machine's global optimum at the person's expense, although the machine never estimated the person's utility (Chasnov et al., 2025).

What this means for PBO. Adaptation drift has structure: variability rises and then falls, parameters adapt at different rates, and novices' preferred magnitudes rise with exposure. A temporal kernel with a fast and a slow lengthscale, in the spirit of the two-state model of motor adaptation (Smith et al., 2006), suits this better than a single exponential decay, and a familiarization phase with lower weights on early comparisons is a simpler alternative (inference). Because the optimizer's update rule changes where the user ends up, the final setting should be compared with alternatives again after a washout period, with the optimizer frozen (inference). Once the remaining differences fall below a parameter's just-noticeable difference, further comparisons carry almost no information, which gives a grounded stopping rule for that parameter; this is where a Weber-style floor does apply (inference). The preferred optimum differs from the physiological one, which supports a multi-objective formulation (Section 14.5) that keeps preference and physiology separate (inference). And simpler methods may suffice: in a study of user-driven manual tuning, people tuned exoskeleton assistance in about 11 minutes, testing 30.5 settings, and reduced metabolic cost by 16.6% (Schäfer et al., 2026). The exoskeleton case study (Chapter 24) and Section 33.1 work through these trade-offs.

Sources cited in Section 39.7 9
  1. Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
  2. Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
  3. Abram et al. (2022) General variability leads to specific adaptation toward optimal movement policies
  4. Clites et al. (2021) Understanding patient preference in prosthetic ankle stiffness
  5. Shepherd et al. (2018) Amputee perception of prosthetic ankle stiffness during locomotion
  6. Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
  7. Chasnov et al. (2025) Human adaptation to adaptive machines converges to game-theoretic equilibria
  8. Smith et al. (2006) Interacting Adaptive Processes with Different Timescales Underlie Short-Term Motor Learning
  9. Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking

39.8 What the evidence changes #

Table 39.1 collects the claims examined in this chapter and where the evidence leaves them; the status column describes only the direction of the evidence. The claims of all chapters in this part, with what each means for PBO, are gathered in Table 45.1 and Table 45.2.

Table 39.1 Claims about preference examined in this chapter, and where the evidence from 2017 to September 2026 leaves them.
Claim Status Key evidence Where
Discrimination declines near the optimum (Weber) fails for value judgments, holds for perceptual parameters high-value decisions faster and more accurate; just-noticeable stiffness difference while walking about 42% Section 39.4, Section 39.7
Attention raises value; unattended discount a constant 0.3 direction holds, effect small; discount varies with overall value and person PP about .53 to .54 Section 39.3.2
Response times carry preference strength strengthened; entered preference learning with guarantees a loss with guarantees; Gaussian process likelihood; overall-value confound Section 39.3.3
Efficient coding explains value noise new and strengthened food values; number coding; risky choice Section 39.4
Ventral striatum predicts song sales small sample, unreplicated; a similar 97% result failed an audit 27 adolescents; 24 songs Section 39.2
Wanting and liking separate in human reports weak support drugs move both or neither Section 39.6
Active inference derives exploration and is a usable alternative to PBO not as of September 2026 the weight is still tuned; no human pairwise experiments Section 39.5
Users adapt during device tuning strengthened, quantified about 109 minutes to expertise; co-adaptation Section 39.7

39.9 Settled, contested, missing #

Research status Settled, contested, missing

Settled. Value signals in OFC and vmPFC are causally related to choice within a task (Ballesta et al., 2020; Lopez-Persem et al., 2020). Attention causally affects choice, but only by a few percentage points (Bhatnagar and Orquin, 2022). Decisions between high-value options are faster and more accurate (Shevlin et al., 2022), so the Weber argument fails for value judgments, while perceptual parameters of devices have measurable discrimination floors (Maberry and Martin, 2026). Response times can enter preference learning with guarantees (Sawarni et al., 2025). Neural forecasts work better for aggregates than for individuals (Hakim et al., 2021). People adapt to a device on several time scales while it is tuned (Poggensee and Collins, 2021).

Contested. Whether a cardinal common currency exists across tasks (Hayden and Niv, 2021). Divisive against range normalization, and which applies to which task. Whether active inference adds anything to PBO beyond a weighted information-gain acquisition function, and whether its foundations hold.

Missing. A PBO study that uses a response-time-augmented Gaussian process with people in the loop, or that models overall value in a response-time likelihood. An attentional DDM fitted to pairwise judgments of continuous design parameters. Any test of common currency or value construction on design parameters. A test of efficient-coding predictions in an interactive optimization session. An expected-free-energy acquisition function compared with EUBO on human pairwise data. Pairwise preference elicitation under pharmacological manipulation. A validation of a tuned device setting after washout with the optimizer frozen.

Sources cited in Section 39.9 9
  1. Ballesta et al. (2020) Values encoded in orbitofrontal cortex are causally related to economic choices
  2. Lopez-Persem et al. (2020) Four core properties of the human brain valuation system demonstrated in intracranial signals
  3. Bhatnagar and Orquin (2022) A meta-analysis on the effect of visual attention on choice
  4. Shevlin et al. (2022) High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity
  5. Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
  6. Sawarni et al. (2025) Preference Learning with Response Time: Robust Losses and Guarantees
  7. Hakim et al. (2021) Machines learn neuromarketing: Improving preference prediction from self-reports using multiple EEG measures and machine learning
  8. Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
  9. Hayden and Niv (2021) The case against economic values in the orbitofrontal cortex (or anywhere else in the brain)

Further reading #

References

  1. Abram, S. J., Poggensee, K. L., Sánchez, N., Simha, S. N., Finley, J. M., Collins, S. H., and Donelan, J. M. (2022). General variability leads to specific adaptation toward optimal movement policies. Current Biology. Cited in §39.7
  2. Ballesta, S., Shi, W., Conen, K. E., and Padoa-Schioppa, C. (2020). Values encoded in orbitofrontal cortex are causally related to economic choices. Nature. Cited in §39.1 §39.9
  3. Bartra, O., McGuire, J. T., and Kable, J. W. (2013). The valuation system: A coordinate-based meta-analysis of BOLD fMRI experiments examining neural correlates of subjective value. NeuroImage. Cited in §39.1
  4. Benkert, J.-M., Liu, S., and Netzer, N. (2026). Time is Knowledge: What Response Times Reveal. working paper (arXiv). working paper Cited in §39.3
  5. Berns, G. S., and Moore, S. E. (2012). A neural predictor of cultural popularity. Journal of Consumer Psychology. doi:10.1016/j.jcps.2011.05.001. Cited in §39.2
  6. Berridge, K. C., and Robinson, T. E. (1998). What is the role of dopamine in reward: hedonic impact, reward learning, or incentive salience? Brain Research Reviews. Cited in §39.6
  7. Bhatnagar, R., and Orquin, J. L. (2022). A meta-analysis on the effect of visual attention on choice. Journal of Experimental Psychology: General. Cited in §39.3 §39.9
  8. Bogacz, R., Brown, E., Moehlis, J., Holmes, P., and Cohen, J. D. (2006). The physics of optimal decision making: A formal analysis of models of performance in two-alternative forced-choice tasks. Psychological Review. Cited in §39.3
  9. Bruineberg, J., Dołęga, K., Dewhurst, J., and Baltieri, M. (2021). The Emperor's New Markov Blankets. Behavioral and Brain Sciences. Cited in §39.5
  10. Busemeyer, J. R., and Townsend, J. T. (1993). Decision field theory: A dynamic-cognitive approach to decision making in an uncertain environment. Psychological Review. Cited in §39.3
  11. Chakroun, K., Mathar, D., Wiehler, A., Ganzer, F., and Peters, J. (2020). Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making. eLife. Cited in §39.6
  12. Chasnov, B. J., Ratliff, L. J., and Burden, S. A. (2025). Human adaptation to adaptive machines converges to game-theoretic equilibria. Scientific Reports. Cited in §39.7
  13. Clites, T. R., Shepherd, M. K., Ingraham, K. A., Wontorcik, L., and Rouse, E. J. (2021). Understanding patient preference in prosthetic ankle stiffness. Journal of NeuroEngineering and Rehabilitation. Cited in §39.7
  14. Conen, K. E., and Padoa-Schioppa, C. (2019). Partial Adaptation to the Value Range in the Macaque Orbitofrontal Cortex. The Journal of Neuroscience. Cited in §39.1
  15. Ferreri, L., Mas-Herrero, E., Zatorre, R. J., Ripollés, P., Gomez-Andres, A., Alicart, H., … Rodriguez-Fornells, A. (2019). Dopamine modulates the reward experiences elicited by music. Proceedings of the National Academy of Sciences. Cited in §39.6
  16. Friston, K. (2010). The free-energy principle: a unified brain theory? Nature Reviews Neuroscience. Cited in §39.5
  17. Friston, K., Rigoli, F., Ognibene, D., Mathys, C., Fitzgerald, T., and Pezzulo, G. (2015). Active inference and epistemic value. Cognitive Neuroscience. Cited in §39.5
  18. Frydman, C., and Jin, L. J. (2022). Efficient Coding and Risky Choice. The Quarterly Journal of Economics. Cited in §39.4
  19. Genevsky, A., Yoon, C., and Knutson, B. (2017). When Brain Beats Behavior: Neuroforecasting Crowdfunding Outcomes. The Journal of Neuroscience. Cited in §39.2
  20. Genevsky, A., Tong, L. C., and Knutson, B. (2025). Neuroforecasting reveals generalizable components of choice. PNAS Nexus. Cited in §39.2
  21. Hakim, A., Klorfeld, S., Sela, T., Friedman, D., Shabat-Simon, M., and Levy, D. J. (2021). Machines learn neuromarketing: Improving preference prediction from self-reports using multiple EEG measures and machine learning. International Journal of Research in Marketing. Cited in §39.2 §39.9
  22. Hayden, B. Y., and Niv, Y. (2021). The case against economic values in the orbitofrontal cortex (or anywhere else in the brain). Behavioral Neuroscience. Cited in §39.1 §39.9
  23. Ingraham, K. A., Remy, C. D., and Rouse, E. J. (2022). The role of user preference in the customized control of robotic exoskeletons. Science Robotics. Cited in §39.7
  24. Kapoor, S., and Narayanan, A. (2023). No, you still cannot predict hit songs using machine learning. Princeton Reproducibility. non-peer-reviewed Cited in §39.2
  25. Kaufmann, T., Metz, Y., Keim, D., and Hüllermeier, E. (2025). ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning. NeurIPS. Cited in §39.3
  26. Kelly, M. A., Patel, R., Thomas, A., Zhu, Z., Quan, Z., Carlson, T., and Cho, Y. (2026). BOBA: Dynamic Bayesian Optimization through Bayesian Active Inference. arXiv preprint 2609.26021. preprint Cited in §39.5
  27. Klein, C. (2018). What do predictive coders want? Synthese. Cited in §39.5
  28. Korb, S., Götzendorfer, S. J., Massaccesi, C., Sezen, P., Graf, I., Willeit, M., Eisenegger, C., and Silani, G. (2020). Dopaminergic and opioidergic regulation during anticipation and consumption of social and nonsocial rewards. eLife. Cited in §39.6
  29. Krajbich, I., Armel, C., and Rangel, A. (2010). Visual fixations and the computation and comparison of value in simple choice. Nature Neuroscience. Cited in §39.3
  30. Lee, D. G., and Daunizeau, J. (2021). Trading mental effort for confidence in the metacognitive control of value-based decision-making. eLife. Cited in §39.1
  31. Levy, D. J., and Glimcher, P. W. (2012). The root of all value: a neural common currency for choice. Current Opinion in Neurobiology. Cited in §39.1
  32. Li, S., Zhang, Y., Ren, Z., Liang, C., Li, N., and Shah, J. A. (2024a). Enhancing Preference-based Linear Bandits via Human Response Time. Advances in Neural Information Processing Systems. Cited in §39.3
  33. Li, Y., Parashar, A., Zhou, E., and Fan, C. (2026c). Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference. arXiv preprint 2602.06029. preprint Cited in §39.5
  34. Li, Y., Parashar, A., Zhou, E., and Fan, C. (2026e). Pragmatic Curiosity: A Unified Framework for Hybrid Learning and Optimization via Active Inference. arXiv preprint 2602.06104. preprint Cited in §39.5
  35. Lopez-Persem, A., Bastin, J., Petton, M., Abitbol, R., Lehongre, K., Adam, C., … Pessiglione, M. (2020). Four core properties of the human brain valuation system demonstrated in intracranial signals. Nature Neuroscience. Cited in §39.1 §39.9
  36. Maberry, A., and Martin, A. E. (2026). Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton. IEEE Transactions on Neural Systems and Rehabilitation Engineering. Cited in §39.7 §39.9
  37. Matějka, F., and McKay, A. (2015). Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model. American Economic Review. Cited in §39.4
  38. Merritt, S. H., Gaffuri, K., and Zak, P. J. (2023). Accurately predicting hit songs using neurophysiology and machine learning. Frontiers in Artificial Intelligence. Cited in §39.2
  39. Millidge, B., Tschantz, A., and Buckley, C. L. (2021). Whence the Expected Free Energy? Neural Computation. Cited in §39.5
  40. Poggensee, K. L., and Collins, S. H. (2021). How adaptation, training, and customization contribute to benefits from exoskeleton assistance. Science Robotics. Cited in §39.7 §39.9
  41. Polanía, R., Woodford, M., and Ruff, C. C. (2019). Efficient coding of subjective value. Nature Neuroscience. Cited in §39.4
  42. Prat-Carrabin, A., and Woodford, M. (2022). Efficient coding of numbers explains decision bias and noise. Nature Human Behaviour. Cited in §39.4
  43. Ratcliff, R. (1978). A theory of memory retrieval. Psychological Review. Cited in §39.3
  44. Sajid, N., Ball, P. J., Parr, T., and Friston, K. J. (2021). Active Inference: Demystified and Compared. Neural Computation. Cited in §39.5
  45. Sawarni, A., Sarmasarkar, S., and Syrgkanis, V. (2025). Preference Learning with Response Time: Robust Losses and Guarantees. NeurIPS. Cited in §39.3 §39.9
  46. Schäfer, N., Zhao, G., Li, B., Kupnik, M., Seyfarth, A., Beckerle, P., and Grimmer, M. (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §39.7
  47. Shen, B., Nguyen, D., Wilson, J., Glimcher, P. W., and Louie, K. (2025b). Early versus late noise differentially enhances or degrades context-dependent choice. Nature Communications. doi:10.1038/s41467-025-59140-3. Cited in §39.4
  48. Shepherd, M. K., Azocar, A. F., Major, M. J., and Rouse, E. J. (2018). Amputee perception of prosthetic ankle stiffness during locomotion. Journal of NeuroEngineering and Rehabilitation. Cited in §39.7
  49. Shevlin, B. R. K., Smith, S. M., Hausfeld, J., and Krajbich, I. (2022). High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity. Proceedings of the National Academy of Sciences. Cited in §39.3 §39.9
  50. Shvartsman, M., Letham, B., Bakshy, E., and Keeley, S. (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Cited in §39.3
  51. Sims, C. A. (2003). Implications of rational inattention. Journal of Monetary Economics. Cited in §39.4
  52. Smith, M. A., Ghazizadeh, A., and Shadmehr, R. (2006). Interacting Adaptive Processes with Different Timescales Underlie Short-Term Motor Learning. PLoS Biology. Cited in §39.7
  53. Sun, H., Shen, Y., and Ton, J.-F. (2025). Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives. ICLR 2025. Cited in §39.1
  54. Ting, C.-C., and Gluth, S. (2025). High overall values mitigate gaze-related effects in perceptual and preferential choices. Journal of Experimental Psychology: General. Cited in §39.3
  55. Wagenmakers, E.-J., Van Der Maas, H. L. J., and Grasman, R. P. P. P. (2007). An EZ-diffusion model for response time and accuracy. Psychonomic Bulletin & Review. Cited in §39.3
  56. Webb, R., Glimcher, P. W., and Louie, K. (2021). The Normalization of Consumer Valuations: Context-Dependent Preferences from Neurobiological Constraints. Management Science. Cited in §39.4
  57. Weintraub, D., Koester, J., Potenza, M. N., Siderowf, A. D., Stacy, M., Voon, V., … Lang, A. E. (2010). Impulse Control Disorders in Parkinson Disease: A Cross-Sectional Study of 3090 Patients. Archives of Neurology. Cited in §39.6
  58. Xing, W. (2026). Attention Limited Reward Learning. arXiv preprint 2607.04590. preprint Cited in §39.3
  59. Zimmermann, J., Glimcher, P. W., and Louie, K. (2018). Multiple timescales of normalized value coding underlie adaptive choice behavior. Nature Communications. Cited in §39.4