Neuroscience and Computational Cognitive Science
The two previous chapters reported behavior: what people choose, how consistently, and under which conditions (Chapter 37, Chapter 38). This chapter asks what the brain and computational models of the mind add for someone who builds or evaluates a preferential Bayesian optimization (PBO) system: whether a single scale of value exists inside the head, as the latent utility of PBO presumes; what a response time reveals beyond the answer; and where the noise in a comparison comes from, and whether it stays fixed.
The main change between 2017 and 2026 was to treat value as constructed during a decision and revised by the choice itself. A scalar value has causal support within one task, but a cardinal "common currency" across tasks remains disputed. The noise in value judgments depends on the values seen before, so discrimination does not decline near the optimum for value judgments, only for perceptual parameters. Attention moves choice causally, but a little. Response times entered preference learning with guarantees, though not yet in live experiments with people. Neural signals add only about 4 percentage points over self-report for an individual, and active inference is not yet a usable alternative to PBO. The centerpiece is the drift-diffusion model of Section 39.3, which you can run in Figure 39.1.
39.1 Neuroeconomics and value-based decisions #
Neuroeconomics looks for the brain's representation of value. A common claim holds that the ventromedial prefrontal cortex (vmPFC) and orbitofrontal cortex (OFC), regions behind the forehead and above the eyes, integrate the values of different kinds of reward into a common currency, a single scale on which food, money, and music can be compared (Levy and Glimcher, 2012; Bartra et al., 2013). A critique, published in Behavioral Neuroscience, argues that value-related signals may reflect salience, arousal, or attention, that OFC coding adapts to the range of values on offer, and that choice can bypass value altogether through direct learning of action policies; its central claim is that a cardinal common-currency value is not, by default, represented in the brain and used for choice (Hayden and Niv, 2021).
Within a task, the causal evidence grew stronger. Electrically stimulating the OFC of macaque monkeys shifted their choices by raising the value of an individual offer (Ballesta et al., 2020), and in intracranial recordings from 36 patients with epilepsy, subjective value could be decoded from vmPFC and lateral OFC activity, with both a linear signal (value) and a quadratic one (confidence) (Lopez-Persem et al., 2020). OFC neurons adapt to the range of values on offer, but only partially, and in a linear decision model the resulting raised floor of activity increases choice variability (Conen and Padoa-Schioppa, 2019). Value is also constructed during the decision. A metacognitive control model, in which a fast, uncertain initial estimate of value is refined under a trade-off between mental effort and confidence, predicts at once the response time, confidence, changes of mind, and choice-induced preference change, and its authors argue that a preference reversal can be a re-evaluation rather than noise (Lee and Daunizeau, 2021). The size of choice-induced preference change is in Section 37.2.2.
What this means for PBO. The evidence supports a scalar latent utility within one task and one session, identifiable only up to a monotone transformation (Section 18.4), in line with the argument that preference learning needs only order consistency (Sun et al., 2025); it does not support reusing a learned scale in another task without recalibration (inference). Because coding partly adapts to the range presented, overly wide exploratory queries early on may leave higher variability later (inference). Confidence is encoded together with value, so scaling the noise of the probit link by the confidence of each comparison is a direct way to implement heteroscedastic noise (inference). All of this evidence comes from food, consumer goods, juice offers, and money, so for continuous design parameters these are working hypotheses (Section 39.9).
Sources cited in Section 39.1 8
- Levy and Glimcher (2012) The root of all value: a neural common currency for choice
- Bartra et al. (2013) The valuation system: A coordinate-based meta-analysis of BOLD fMRI experiments examining neural correlates of subjective value
- Hayden and Niv (2021) The case against economic values in the orbitofrontal cortex (or anywhere else in the brain)
- Ballesta et al. (2020) Values encoded in orbitofrontal cortex are causally related to economic choices
- Lopez-Persem et al. (2020) Four core properties of the human brain valuation system demonstrated in intracranial signals
- Conen and Padoa-Schioppa (2019) Partial Adaptation to the Value Range in the Macaque Orbitofrontal Cortex
- Lee and Daunizeau (2021) Trading mental effort for confidence in the metacognitive control of value-based decision-making
- Sun et al. (2025) Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
39.2 Neuromarketing and neuroforecasting #
A common claim holds that neural signals reveal preferences people cannot report. Its best-known support is a study in which activity in the nucleus accumbens (part of the ventral striatum) of adolescents listening to songs predicted the songs' later sales while their own ratings did not (Berns and Moore, 2012). The evidence base is small: after exclusions there were 27 adolescents; 87 of the 120 songs had sales data; accumbens activity correlated with log sales at (), the mean likability rating at (not significant); a logistic regression classified only 30% of the hits correctly; and the data came from an experiment originally designed to study social influence (Berns and Moore, 2012). We found no independent direct replication. A later paper reporting 97% accuracy in classifying hit songs (Merritt et al., 2023) was audited: the data had been oversampled before they were split into training and test sets, covered only 24 songs and 33 listeners, and the corrected accuracy was hardly better than chance (Kapoor and Narayanan, 2023), a non-peer-reviewed audit.
The sturdier progress is at the level of populations. In 30 participants scanned with functional magnetic resonance imaging (fMRI), nucleus accumbens activity predicted the outcomes of internet crowdfunding campaigns weeks later while the participants' own choices did not, and the result replicated in a second study (Genevsky et al., 2017). With internet samples of 2,956 and 992 people, forecasts based on behavior varied with the demographic representativeness of the laboratory sample, while forecasts based on brain activity stayed significant (Genevsky et al., 2025). For individuals the gain is limited: combining several electroencephalography (EEG) measures with machine learning predicted each person's most and least liked of six food products with 68.5% accuracy, 4.07 percentage points above self-report alone (Hakim et al., 2021). We found no meta-analysis of the accuracy of neural prediction, and the out-of-sample evidence comes mainly from one laboratory.
What this means for PBO. PBO learns one person's utility, while the advantage of neural forecasting is documented for aggregates, so its most defensible use is a population prior or a warm start, not a personal likelihood (inference). A session of a few dozen comparisons cannot train a personal decoder, and any PBO system that uses a neural proxy needs strict held-out evaluation (inference).
Sources cited in Section 39.2 6
- Berns and Moore (2012) A neural predictor of cultural popularity
- Merritt et al. (2023) Accurately predicting hit songs using neurophysiology and machine learning
- Kapoor and Narayanan (2023) No, you still cannot predict hit songs using machine learning
- Genevsky et al. (2017) When Brain Beats Behavior: Neuroforecasting Crowdfunding Outcomes
- Genevsky et al. (2025) Neuroforecasting reveals generalizable components of choice
- Hakim et al. (2021) Machines learn neuromarketing: Improving preference prediction from self-reports using multiple EEG measures and machine learning
39.3 Sequential sampling and response times #
A common claim, drawing on decision field theory (Busemeyer and Townsend, 1993) and the attentional drift-diffusion model (Krajbich et al., 2010), treats response time as a natural measure of preference uncertainty: a fast answer signals a large difference in value, a slow one a small difference. Two theoretical results back it: in preference-based bandits, response times are most informative for easy queries (Li et al., 2024a), and choices and response times together can identify the distribution of latent preferences at several points, according to a 2026 working paper (Benkert et al., 2026).
39.3.1 The drift-diffusion model #
The drift-diffusion model (DDM) describes a two-option decision as noisy evidence accumulating over time (Ratcliff, 1978). Think of a counter that starts at zero. Every few milliseconds it moves up a little if the sample of evidence favors option A and down if it favors B; on average it drifts at rate , the drift rate, but each step is noisy. The decision is made when the counter first reaches (choose A) or (choose B); the boundary expresses caution. The response time is the time to reach a boundary plus a fixed non-decision time for perceiving and pressing a key. In value-based choice the drift is proportional to the difference in value, with .
Two closed forms describe the model with unit noise (Bogacz et al., 2006). The probability of choosing A and the mean decision time are
The first is the logistic function of . A model of the decision process, in other words, produces the logistic (Bradley-Terry) link of Section 16.4 as its choice probability, with a slope set jointly by the person's caution and the drift scale .
Derivation Why the drift-diffusion model gives the logistic link
- Let be the probability of reaching before when the counter is at . Over a short time the counter moves by plus noise of variance , so with .
- Expand to second order and take the expectation, using and : , so .
- The general solution of this linear equation is , as substituting it confirms.
- The boundaries fix the constants: and give .
- At the start, . Factor the denominator as and cancel: , the logistic function of .
The model makes the claim about response times precise. When the drift vanishes, choices are coin flips, and the mean decision time takes its largest value, . As grows, the choice probability saturates near 0 or 1 while the response time keeps falling. Beyond a moderate utility difference the choice says almost nothing new, but the response time still does, which is why response times help most for easy queries (Li et al., 2024a).
Things to try:
- Set . About half of the simulated decisions end at each boundary, and the decision-time distributions are widest.
- Raise to 1 and beyond. The choice probability is already near 1, while the mean response time keeps dropping: in this range only the response time distinguishes a large utility difference from a very large one.
- Raise the caution . Choices become more accurate and slower, and the logistic curve steepens: the slope of the link, which a comparison model absorbs into its noise scale , reflects the person's caution as well as their utility.
- Switch on the overall-value effect and raise . The same utility difference now yields a faster and more accurate choice, and a model that ignores overall value would read the response as a larger utility gap.
39.3.2 Attention and response times, measured #
The causal effect of attention now has a number, which Figure 39.2 shows. Pooling experiments that manipulated attention in two-option preferential choice, increasing the total time an option was looked at raised the probability of choosing it to (95% confidence interval ), controlling which option was looked at last gave (), and manipulating the first fixation had no effect (, , ); slight publication bias lowers these values a little (Bhatnagar and Orquin, 2022). The discount on the unattended option is not a constant. In the study that introduced the attentional model, 0.3 was the best fit to the pooled data of 39 participants, while fits to each participant gave a mean of 0.52 with a standard deviation of 0.3 (Krajbich et al., 2010); in a preregistered study (), gaze effects weakened when the overall value of the options was high (Ting and Gluth, 2025).
Things to try:
- Compare the three intervals with the 0.5 line. Those for total looking time (.523 to .560) and for the last fixation (.518 to .547) lie above it; the one for the first fixation (.497 to .516) contains it.
- Select total looking time and set the session to 40 comparisons. If every comparison carried that manipulation, the option looked at longer would win about 1.6 more answers (0.9 to 2.4) out of 40 than with no effect.
- In the lower panel, the pooled discount, 0.3, lies below the mean of the individual fits, 0.52, whose standard deviation, 0.3, is as large as the pooled value itself.
Response time measures preference strength, with confounds. In three domains of choice, decisions between two high-value options were mostly faster and more accurate, which in the DDM means a higher drift rate, contrary to diminishing sensitivity to value (Shevlin et al., 2022); the negative relation between overall value and response time replicated in Ting and Gluth (2025). Response-time proxies for preference strength are often noisy and confounded (Kaufmann et al., 2025).
39.3.3 Response times enter preference learning #
The main development of 2017 to 2026 is that response times entered preference learning with guarantees, Gaussian processes included. Shvartsman et al. (2024) approximated the DDM likelihood differentiably with a family of skewed three-parameter distributions, so that response times can enter a Gaussian process model of binary choice, and report better estimates of latent value and predictions of held-out choices on three real psychophysics and preference data sets (UAI 2024; Section 27.2 gives the details). Sawarni et al. (2025) built a loss that uses response times under the EZ-diffusion model, a simplified DDM whose parameters can be computed in closed form from the mean and variance of response times and the accuracy (Wagenmakers et al., 2007); for linear rewards it turns the error of standard preference learning, which grows exponentially with the magnitude of the reward, into polynomial growth, and it extends to nonparametric reward spaces. A July 2026 preprint analyzes each pairwise label under rational inattention, the idea that people pay for information and attend only as much as it is worth (Section 40.5). It argues that heterogeneous attention can lead a Bradley-Terry model to recover a misleading ranking, and that in perceptual comparisons response times and gaze carry information about the size of the gap that labels lack (Xing, 2026).
What this means for PBO. A joint likelihood of choice and response time can already be used with a Gaussian process surrogate. It should take overall value (or a proxy, such as the current posterior means of the two options) as a covariate of the drift rate or the boundary; otherwise fast answers late in a session, when both options are good, will be read as large utility differences (inference), the misreading Figure 39.1 shows. Simultaneous, side-by-side presentation should be preferred, and options that must be experienced in sequence (audio, touch, gait) should have their order counterbalanced, with an order bias term in the likelihood; the expected presentation effect is a few percentage points of choice probability (inference). The spread of individual discounts supports per-user bias or discount parameters rather than a population constant such as 0.3 (inference). The existing response-time Gaussian process work used offline data; the experiments that would test these recommendations with people in the loop are listed in Section 39.9.
Sources cited in Section 39.3 14
- Busemeyer and Townsend (1993) Decision field theory: A dynamic-cognitive approach to decision making in an uncertain environment
- Krajbich et al. (2010) Visual fixations and the computation and comparison of value in simple choice
- Li et al. (2024a) Enhancing Preference-based Linear Bandits via Human Response Time
- Benkert et al. (2026) Time is Knowledge: What Response Times Reveal
- Ratcliff (1978) A theory of memory retrieval
- Bogacz et al. (2006) The physics of optimal decision making: A formal analysis of models of performance in two-alternative forced-choice tasks
- Bhatnagar and Orquin (2022) A meta-analysis on the effect of visual attention on choice
- Ting and Gluth (2025) High overall values mitigate gaze-related effects in perceptual and preferential choices
- Shevlin et al. (2022) High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity
- Kaufmann et al. (2025) ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning
- Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences
- Sawarni et al. (2025) Preference Learning with Response Time: Robust Losses and Guarantees
- Wagenmakers et al. (2007) An EZ-diffusion model for response time and accuracy
- Xing (2026) Attention Limited Reward Learning
39.4 Efficient coding and normalization #
The Weber argument holds that as the optimizer converges on the high-utility region, the person's ability to tell options apart declines, so queries should target regions where the expected utility difference exceeds the just-noticeable difference (the smallest difference a person reliably detects). A related argument treats comparison noise as fixed expression noise around a stable utility, with rational inattention (Sims, 2003; Matějka and McKay, 2015) as its microfoundation.
For value judgments, the Weber argument fails: decisions between high-value options are faster and more accurate (Section 39.3.2). Efficient coding, the principle that a system with limited capacity should spend its resolution where inputs are most frequent, predicts that discrimination tracks the density of values in the environment rather than their magnitude. Polanía et al. (2019) had 38 participants rate and choose among 64 familiar foods, the first of four experiments with 127 participants in all: a valuation model that maximizes information under limited coding resources explained rating variability, choice consistency, and confidence together, with choices predicted to be more consistent at values where the prior density is high. Related work found that rarer numbers are coded more noisily (Prat-Carrabin and Woodford, 2022), that in two preregistered experiments risk-taking was more sensitive to payoffs that occurred more often (Frydman and Jin, 2022), and that monkeys chose more sensitively in environments with lower reward variance (Zimmermann et al., 2018). The noise itself changes with context and history, then, rather than being fixed expression noise around a stable utility.
Divisive normalization, in which each option's value is divided by a measure of the total value on offer, remains contested. A discrete choice model with divisive normalization fitted effects of choice-set size and composition better than alternatives including range normalization (Webb et al., 2021), but a distractor effect attributed to normalization failed to replicate (Section 37.2.3), and in tasks with learned values range normalization won (Section 37.3.1). A distinction between early and late noise offers a reconciliation: when the representation of the options is blurry (early noise), a third option improves discrimination, and when time pressure dominates (late noise), the context harms it; a divisive normalization model with both kinds of noise fitted best (Shen et al., 2025b).
What this means for PBO. An acquisition function that concentrates queries near the incumbent narrows the range of utilities the person experiences. Efficient coding predicts that sensitivity in that region rises during the session, partially and with a lag, and a probit likelihood with a fixed noise scale would misread that late precision as a steeper utility surface (inference). A noise scale that depends on how typical the options' utilities are, relative to those presented so far, can represent this; a simpler practice is to insert a few fixed reference pairs from rarely visited regions and measure discrimination there directly (inference). Whether Weber-style compression or efficient coding dominates in a domain can be checked with a few repeated pairs at different stages of a session: for value judgments the evidence favors efficient coding, while perceptual parameters (stiffness, torque, color) have floors (Section 39.7) (inference).
Sources cited in Section 39.4 8
- Sims (2003) Implications of rational inattention
- Matějka and McKay (2015) Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model
- Polanía et al. (2019) Efficient coding of subjective value
- Prat-Carrabin and Woodford (2022) Efficient coding of numbers explains decision bias and noise
- Frydman and Jin (2022) Efficient Coding and Risky Choice
- Zimmermann et al. (2018) Multiple timescales of normalized value coding underlie adaptive choice behavior
- Webb et al. (2021) The Normalization of Consumer Valuations: Context-Dependent Preferences from Neurobiological Constraints
- Shen et al. (2025b) Early versus late noise differentially enhances or degrades context-dependent choice
39.5 Predictive processing and active inference #
The free-energy principle proposes that brains, and perhaps all self-organizing systems, act to minimize variational free energy, a bound on how surprising their sensations are under their internal model (Friston, 2010). Active inference applies it to action: an agent chooses actions that minimize expected free energy, and preferences are prior beliefs about the outcomes the agent expects to observe. Expected free energy decomposes into a pragmatic value (reaching preferred outcomes) and an epistemic value (gaining information) (Friston et al., 2015). The decomposition has been called "exactly" the balance of exploration and exploitation in PBO, with an exploration term "derived from first principles rather than added by hand" like the weight of an upper confidence bound (Section 12.4).
That claim needs qualification. A functional argued to be the natural extension of variational free energy into the future actively suppresses exploration, so exploration does not follow directly from minimizing free energy (Millidge et al., 2021). The first theoretical guarantee for agents that minimize expected free energy, a 2026 preprint, shows that "sufficient curiosity" ensures both posterior consistency and bounded cumulative regret, while too little curiosity leads to myopic exploitation and too much to unnecessary exploration and regret (Li et al., 2026c). The weight, in other words, still has to be set, just like the exploration weight of an upper confidence bound. The foundations are contested too: Markov blankets, the boundaries that separate a system from its environment in the theory, are persistently conflated between a tool of inference and physical boundaries (Bruineberg et al., 2021), and the "dark room problem", that an agent minimizing surprise should seek a dark, quiet room, is a real difficulty for the idea that preference is prediction (Klein, 2018). Three 2026 preprints combine active inference with Bayesian optimization: the regret guarantee above; "pragmatic curiosity", which scores queries by the information gained about task-relevant latent variables plus a potential of expected regret (Li et al., 2026e); and BOBA, an acquisition function that adds lookahead uncertainty about how a time-varying objective will change, which improved regret on synthetic dynamic benchmarks (Kelly et al., 2026).
What this means for PBO. As of September 2026, active inference is not a usable alternative to PBO (inference, from the studies above). Every concrete implementation in Bayesian optimization reduces to an acquisition function, a pragmatic term plus a weighted information-gain term, close to existing information-theoretic acquisition functions, and it replaces neither the Gaussian process preference model nor the pairwise likelihood. The abstracts of the three preprints report no experiments with human pairwise feedback (we did not read the full texts), and no study compares an expected-free-energy acquisition function with EUBO-type acquisition functions (Section 19.4) on human preference data. And the two frameworks mean different things by preference: in active inference it is the agent's prior over its own observations, in PBO the person's utility is a hidden state to be inferred, so rewriting PBO as active inference re-encodes the objective without adding any constraint on the structure of human preference (inference). One role could be useful: modeling the person as an agent whose prior preferences are learned and updated (Sajid et al., 2021) would give a generative model of preference drift, though we found no human preference data fitted that way (inference).
Sources cited in Section 39.5 9
- Friston (2010) The free-energy principle: a unified brain theory?
- Friston et al. (2015) Active inference and epistemic value
- Millidge et al. (2021) Whence the Expected Free Energy?
- Li et al. (2026c) Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference
- Bruineberg et al. (2021) The Emperor's New Markov Blankets
- Klein (2018) What do predictive coders want?
- Li et al. (2026e) Pragmatic Curiosity: A Unified Framework for Hybrid Learning and Optimization via Active Inference
- Kelly et al. (2026) BOBA: Dynamic Bayesian Optimization through Bayesian Active Inference
- Sajid et al. (2021) Active Inference: Demystified and Compared
39.6 Psychopharmacology and addiction #
Pharmacology offers what look like the strongest counterexamples to stable preference. A common claim distinguishes "wanting", attributed to mesolimbic dopamine, from "liking", attributed to opioid hedonic hotspots in the brain (Berridge and Robinson, 1998), and proposes modeling the two with a multi-output Gaussian process. Evidence that the two separate in human self-report is weak. In a double-blind within-subject study (), levodopa raised and risperidone lowered both musical pleasure and music-related motivation together (Ferreri et al., 2019); in 131 volunteers, the dopamine antagonist amisulpride and the opioid antagonist naltrexone both reduced physical effort, but neither changed subjective ratings of wanting or liking (Korb et al., 2020). Dopaminergic drugs do change behavior. In a cross-sectional study of 3,090 treated patients with Parkinson's disease, impulse control disorders were more common among patients taking a dopamine agonist than among those not taking one (17.1% against 6.9%, odds ratio 2.72) (Weintraub et al., 2010), and in 31 healthy men levodopa weakened directed exploration (Chakroun et al., 2020).
What this means for PBO. Pairwise choices and ratings may not separate wanting from liking, so a two-output model would need different observation channels, such as effort for wanting and ratings during consumption for liking (inference). Medication state is a slow context variable: studies of PBO in health and rehabilitation should record it and allow for drift over weeks that need not be monotone, and since exploration and noise vary with dopamine state, noise parameters should be fixed per session rather than per person (inference). When a drug changes what a person prefers, which preference counts is a question of authority; in clinical settings a reflective endorsement check, made away from the peak effect of a drug or at follow-up, is a reasonable safeguard (inference), one Section 41.2 discusses. No study has elicited pairwise preferences under a pharmacological manipulation.
Sources cited in Section 39.6 5
- Berridge and Robinson (1998) What is the role of dopamine in reward: hedonic impact, reward learning, or incentive salience?
- Ferreri et al. (2019) Dopamine modulates the reward experiences elicited by music
- Korb et al. (2020) Dopaminergic and opioidergic regulation during anticipation and consumption of social and nonsocial rewards
- Weintraub et al. (2010) Impulse Control Disorders in Parkinson Disease: A Cross-Sectional Study of 3090 Patients
- Chakroun et al. (2020) Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making
39.7 Motor learning and adaptation #
A common claim holds that PBO works well where preferences are experiential and users have real expertise. The claim about expertise needs refining: in an ankle exoskeleton, users the study calls knowledgeable (the two groups differed in technical background) preferred higher torque than naive users, and naive users' preferred torque rose over the course of the experiment, so knowledge and exposure change the preferred setting itself, not just the noise around it (Ingraham et al., 2022).
People adapt on several time scales while a device is tuned for them. In Poggensee and Collins (2021), after training with moderate variability, personalized assistance reduced metabolic cost by 39% relative to the exoskeleton switched off, training contributed about half of the benefit and personalization about a quarter, and becoming an expert user took about 109 minutes of assisted walking. When people met a new exoskeleton context, the variability of step frequency, ankle angle range, and muscle activity first rose and then fell, on time scales that differed between variables (Abram et al., 2022). Preferences are repeatable but change: when 12 naive and 12 knowledgeable participants tuned the torque magnitude and timing of an ankle exoskeleton themselves, their preferences ranged from 7.9 to 19.4 N·m and from 54.1% to 59.2% of the gait cycle, with trial-to-trial standard deviations of 1.7 N·m and 1.5%, and convergence within a trial took 105 seconds (Ingraham et al., 2022). The preferred stiffness of a prosthetic ankle maximized the kinematic symmetry between the prosthetic and the intact joint and was not significantly related to body weight or metabolic rate (Clites et al., 2021). Perceptual resolution differs by a factor of about five between judging the stiffness of a prosthetic ankle and that of an ankle exoskeleton (inference, from two studies with different devices and people): eight people with below-knee amputations detected a 7.7% change in stiffness with 75% accuracy (Shepherd et al., 2018), while walking in an ankle exoskeleton the just-noticeable difference was about 42% for stiffness (Maberry and Martin, 2026). The optimizer and the person also adapt to each other: in three online experiments on co-adaptation, when the machine used a policy gradient, human behavior moved toward the machine's global optimum at the person's expense, although the machine never estimated the person's utility (Chasnov et al., 2025).
What this means for PBO. Adaptation drift has structure: variability rises and then falls, parameters adapt at different rates, and novices' preferred magnitudes rise with exposure. A temporal kernel with a fast and a slow lengthscale, in the spirit of the two-state model of motor adaptation (Smith et al., 2006), suits this better than a single exponential decay, and a familiarization phase with lower weights on early comparisons is a simpler alternative (inference). Because the optimizer's update rule changes where the user ends up, the final setting should be compared with alternatives again after a washout period, with the optimizer frozen (inference). Once the remaining differences fall below a parameter's just-noticeable difference, further comparisons carry almost no information, which gives a grounded stopping rule for that parameter; this is where a Weber-style floor does apply (inference). The preferred optimum differs from the physiological one, which supports a multi-objective formulation (Section 14.5) that keeps preference and physiology separate (inference). And simpler methods may suffice: in a study of user-driven manual tuning, people tuned exoskeleton assistance in about 11 minutes, testing 30.5 settings, and reduced metabolic cost by 16.6% (Schäfer et al., 2026). The exoskeleton case study (Chapter 24) and Section 33.1 work through these trade-offs.
Sources cited in Section 39.7 9
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
- Abram et al. (2022) General variability leads to specific adaptation toward optimal movement policies
- Clites et al. (2021) Understanding patient preference in prosthetic ankle stiffness
- Shepherd et al. (2018) Amputee perception of prosthetic ankle stiffness during locomotion
- Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
- Chasnov et al. (2025) Human adaptation to adaptive machines converges to game-theoretic equilibria
- Smith et al. (2006) Interacting Adaptive Processes with Different Timescales Underlie Short-Term Motor Learning
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
39.8 What the evidence changes #
Table 39.1 collects the claims examined in this chapter and where the evidence leaves them; the status column describes only the direction of the evidence. The claims of all chapters in this part, with what each means for PBO, are gathered in Table 45.1 and Table 45.2.
| Claim | Status | Key evidence | Where |
|---|---|---|---|
| Discrimination declines near the optimum (Weber) | fails for value judgments, holds for perceptual parameters | high-value decisions faster and more accurate; just-noticeable stiffness difference while walking about 42% | Section 39.4, Section 39.7 |
| Attention raises value; unattended discount a constant 0.3 | direction holds, effect small; discount varies with overall value and person | about .53 to .54 | Section 39.3.2 |
| Response times carry preference strength | strengthened; entered preference learning with guarantees | a loss with guarantees; Gaussian process likelihood; overall-value confound | Section 39.3.3 |
| Efficient coding explains value noise | new and strengthened | food values; number coding; risky choice | Section 39.4 |
| Ventral striatum predicts song sales | small sample, unreplicated; a similar 97% result failed an audit | 27 adolescents; 24 songs | Section 39.2 |
| Wanting and liking separate in human reports | weak support | drugs move both or neither | Section 39.6 |
| Active inference derives exploration and is a usable alternative to PBO | not as of September 2026 | the weight is still tuned; no human pairwise experiments | Section 39.5 |
| Users adapt during device tuning | strengthened, quantified | about 109 minutes to expertise; co-adaptation | Section 39.7 |
39.9 Settled, contested, missing #
Settled. Value signals in OFC and vmPFC are causally related to choice within a task (Ballesta et al., 2020; Lopez-Persem et al., 2020). Attention causally affects choice, but only by a few percentage points (Bhatnagar and Orquin, 2022). Decisions between high-value options are faster and more accurate (Shevlin et al., 2022), so the Weber argument fails for value judgments, while perceptual parameters of devices have measurable discrimination floors (Maberry and Martin, 2026). Response times can enter preference learning with guarantees (Sawarni et al., 2025). Neural forecasts work better for aggregates than for individuals (Hakim et al., 2021). People adapt to a device on several time scales while it is tuned (Poggensee and Collins, 2021).
Contested. Whether a cardinal common currency exists across tasks (Hayden and Niv, 2021). Divisive against range normalization, and which applies to which task. Whether active inference adds anything to PBO beyond a weighted information-gain acquisition function, and whether its foundations hold.
Missing. A PBO study that uses a response-time-augmented Gaussian process with people in the loop, or that models overall value in a response-time likelihood. An attentional DDM fitted to pairwise judgments of continuous design parameters. Any test of common currency or value construction on design parameters. A test of efficient-coding predictions in an interactive optimization session. An expected-free-energy acquisition function compared with EUBO on human pairwise data. Pairwise preference elicitation under pharmacological manipulation. A validation of a tuned device setting after washout with the optimizer frozen.
Sources cited in Section 39.9 9
- Ballesta et al. (2020) Values encoded in orbitofrontal cortex are causally related to economic choices
- Lopez-Persem et al. (2020) Four core properties of the human brain valuation system demonstrated in intracranial signals
- Bhatnagar and Orquin (2022) A meta-analysis on the effect of visual attention on choice
- Shevlin et al. (2022) High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity
- Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
- Sawarni et al. (2025) Preference Learning with Response Time: Robust Losses and Guarantees
- Hakim et al. (2021) Machines learn neuromarketing: Improving preference prediction from self-reports using multiple EEG measures and machine learning
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
- Hayden and Niv (2021) The case against economic values in the orbitofrontal cortex (or anywhere else in the brain)
Further reading #
- Bogacz et al. (2006) derives the drift-diffusion model's closed forms and relates them to optimal decision making; it is the clearest route from the model to Equation (39.1).
- Shvartsman et al. (2024) and Sawarni et al. (2025) are the two most direct bridges from response times to preference learning, one with Gaussian processes and one with guarantees.
- Shevlin et al. (2022) is a short paper whose result, that high-value decisions are fast and accurate, changes how late-session comparisons should be read.
- Bhatnagar and Orquin (2022) puts a number on how much attention moves choice.
- Polanía et al. (2019) shows efficient coding of subjective value with ordinary foods, and is a good entry to noise that depends on context.
- Millidge et al. (2021) and Li et al. (2026c) together explain why the exploration term of active inference still needs a weight.
- Poggensee and Collins (2021) and Ingraham et al. (2022) are the best quantitative accounts of how people change while a device is tuned for them.
References
- (2022). General variability leads to specific adaptation toward optimal movement policies. Current Biology. Cited in §39.7
- (2020). Values encoded in orbitofrontal cortex are causally related to economic choices. Nature. Cited in §39.1 §39.9
- (2013). The valuation system: A coordinate-based meta-analysis of BOLD fMRI experiments examining neural correlates of subjective value. NeuroImage. Cited in §39.1
- (2026). Time is Knowledge: What Response Times Reveal. working paper (arXiv). working paper Cited in §39.3
- (2012). A neural predictor of cultural popularity. Journal of Consumer Psychology. doi:10.1016/j.jcps.2011.05.001. Cited in §39.2
- (1998). What is the role of dopamine in reward: hedonic impact, reward learning, or incentive salience? Brain Research Reviews. Cited in §39.6
- (2022). A meta-analysis on the effect of visual attention on choice. Journal of Experimental Psychology: General. Cited in §39.3 §39.9
- (2006). The physics of optimal decision making: A formal analysis of models of performance in two-alternative forced-choice tasks. Psychological Review. Cited in §39.3
- (2021). The Emperor's New Markov Blankets. Behavioral and Brain Sciences. Cited in §39.5
- (1993). Decision field theory: A dynamic-cognitive approach to decision making in an uncertain environment. Psychological Review. Cited in §39.3
- (2020). Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making. eLife. Cited in §39.6
- (2025). Human adaptation to adaptive machines converges to game-theoretic equilibria. Scientific Reports. Cited in §39.7
- (2021). Understanding patient preference in prosthetic ankle stiffness. Journal of NeuroEngineering and Rehabilitation. Cited in §39.7
- (2019). Partial Adaptation to the Value Range in the Macaque Orbitofrontal Cortex. The Journal of Neuroscience. Cited in §39.1
- (2019). Dopamine modulates the reward experiences elicited by music. Proceedings of the National Academy of Sciences. Cited in §39.6
- (2010). The free-energy principle: a unified brain theory? Nature Reviews Neuroscience. Cited in §39.5
- (2015). Active inference and epistemic value. Cognitive Neuroscience. Cited in §39.5
- (2022). Efficient Coding and Risky Choice. The Quarterly Journal of Economics. Cited in §39.4
- (2017). When Brain Beats Behavior: Neuroforecasting Crowdfunding Outcomes. The Journal of Neuroscience. Cited in §39.2
- (2025). Neuroforecasting reveals generalizable components of choice. PNAS Nexus. Cited in §39.2
- (2021). Machines learn neuromarketing: Improving preference prediction from self-reports using multiple EEG measures and machine learning. International Journal of Research in Marketing. Cited in §39.2 §39.9
- (2021). The case against economic values in the orbitofrontal cortex (or anywhere else in the brain). Behavioral Neuroscience. Cited in §39.1 §39.9
- (2022). The role of user preference in the customized control of robotic exoskeletons. Science Robotics. Cited in §39.7
- (2023). No, you still cannot predict hit songs using machine learning. Princeton Reproducibility. non-peer-reviewed Cited in §39.2
- (2025). ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning. NeurIPS. Cited in §39.3
- (2026). BOBA: Dynamic Bayesian Optimization through Bayesian Active Inference. arXiv preprint 2609.26021. preprint Cited in §39.5
- (2018). What do predictive coders want? Synthese. Cited in §39.5
- (2020). Dopaminergic and opioidergic regulation during anticipation and consumption of social and nonsocial rewards. eLife. Cited in §39.6
- (2010). Visual fixations and the computation and comparison of value in simple choice. Nature Neuroscience. Cited in §39.3
- (2021). Trading mental effort for confidence in the metacognitive control of value-based decision-making. eLife. Cited in §39.1
- (2012). The root of all value: a neural common currency for choice. Current Opinion in Neurobiology. Cited in §39.1
- (2024a). Enhancing Preference-based Linear Bandits via Human Response Time. Advances in Neural Information Processing Systems. Cited in §39.3
- (2026c). Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference. arXiv preprint 2602.06029. preprint Cited in §39.5
- (2026e). Pragmatic Curiosity: A Unified Framework for Hybrid Learning and Optimization via Active Inference. arXiv preprint 2602.06104. preprint Cited in §39.5
- (2020). Four core properties of the human brain valuation system demonstrated in intracranial signals. Nature Neuroscience. Cited in §39.1 §39.9
- (2026). Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton. IEEE Transactions on Neural Systems and Rehabilitation Engineering. Cited in §39.7 §39.9
- (2015). Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model. American Economic Review. Cited in §39.4
- (2023). Accurately predicting hit songs using neurophysiology and machine learning. Frontiers in Artificial Intelligence. Cited in §39.2
- (2021). Whence the Expected Free Energy? Neural Computation. Cited in §39.5
- (2021). How adaptation, training, and customization contribute to benefits from exoskeleton assistance. Science Robotics. Cited in §39.7 §39.9
- (2019). Efficient coding of subjective value. Nature Neuroscience. Cited in §39.4
- (2022). Efficient coding of numbers explains decision bias and noise. Nature Human Behaviour. Cited in §39.4
- (1978). A theory of memory retrieval. Psychological Review. Cited in §39.3
- (2021). Active Inference: Demystified and Compared. Neural Computation. Cited in §39.5
- (2025). Preference Learning with Response Time: Robust Losses and Guarantees. NeurIPS. Cited in §39.3 §39.9
- (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §39.7
- (2025b). Early versus late noise differentially enhances or degrades context-dependent choice. Nature Communications. doi:10.1038/s41467-025-59140-3. Cited in §39.4
- (2018). Amputee perception of prosthetic ankle stiffness during locomotion. Journal of NeuroEngineering and Rehabilitation. Cited in §39.7
- (2022). High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity. Proceedings of the National Academy of Sciences. Cited in §39.3 §39.9
- (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Cited in §39.3
- (2003). Implications of rational inattention. Journal of Monetary Economics. Cited in §39.4
- (2006). Interacting Adaptive Processes with Different Timescales Underlie Short-Term Motor Learning. PLoS Biology. Cited in §39.7
- (2025). Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives. ICLR 2025. Cited in §39.1
- (2025). High overall values mitigate gaze-related effects in perceptual and preferential choices. Journal of Experimental Psychology: General. Cited in §39.3
- (2007). An EZ-diffusion model for response time and accuracy. Psychonomic Bulletin & Review. Cited in §39.3
- (2021). The Normalization of Consumer Valuations: Context-Dependent Preferences from Neurobiological Constraints. Management Science. Cited in §39.4
- (2010). Impulse Control Disorders in Parkinson Disease: A Cross-Sectional Study of 3090 Patients. Archives of Neurology. Cited in §39.6
- (2026). Attention Limited Reward Learning. arXiv preprint 2607.04590. preprint Cited in §39.3
- (2018). Multiple timescales of normalized value coding underlie adaptive choice behavior. Nature Communications. Cited in §39.4