Judgment, Decision, and Psychophysics
Part IV built preferential Bayesian optimization (PBO) on one modeling decision: a person's answer to "which do you prefer?" is a noisy reading of a fixed latent utility. In Section 27.1, the probability that beats is , no matter when the question is asked, what came before it, or what else is on the screen. Part IX asks whether a preference of that kind is there to be found. This chapter begins with the disciplines that study the act of comparing itself: judgment and decision-making research, psychophysics, mathematical psychology, and the study of heuristics.
Their evidence from 2017 to September 2026 points two ways. Findings once used to argue that people tire quickly shrank to near zero in large preregistered tests, while effects that break the independence of successive comparisons were confirmed and now come with mechanisms a likelihood could include. The picture is a mixed observation model: a stable component, evaluation noise with structure, and drift caused by the queries themselves (inference).
37.1 What the comparison model assumes #
The likelihood of Section 27.1 makes more commitments than its formula shows. Multiplying the likelihoods of separate comparisons, as every PBO paper does, treats answers as conditionally independent given the utility. Four assumptions follow, and this chapter tests each: stability, one utility for the whole session (Section 37.2.2, Section 37.3.1); order independence, no carryover from earlier questions (Section 37.3.2, Section 37.4.3); constant noise, one noise scale for every pair (Section 37.3.1, Section 37.4.1); and context independence, no influence from other options on screen or in memory (Section 37.2.3).
The evidence is weighed with psychology's own yardsticks. The most common effect size, Cohen's , is the difference between two group means divided by the standard deviation within the groups; Cohen's conventions call 0.2 small, 0.5 medium, and 0.8 large (Cohen, 1988), and at the two distributions overlap almost completely. A 95% confidence interval is the range of effect sizes compatible with the data. A meta-analysis pools many studies, and publication bias, the tendency of significant results to be published more often, inflates its estimate. Preregistration fixes the hypotheses and analysis before data collection; a multi-lab replication runs one protocol in many laboratories, and a registered replication report is one accepted for publication before its results are known. A Bayes factor of 4 in favor of the null means the data are four times as probable if there is no effect.
Test-retest reliability, the correlation between two measurements of the same quantity on the same people, decides whether a per-user parameter can be estimated at all. A meta-analysis of the stability of risk preference, built on test-retest correlations, estimated a reliability of 0.61 for self-reported risk propensity and of 0.25 for behavioral measures, tasks such as lottery choices and balloon-pumping games played for real or hypothetical money (Bagaïni et al., 2025): the "behavioral" measure of a stable trait is the less stable one, and at 0.25 most of the variation between people in a single measurement is noise. For calibration, the Many Labs 2 project ran 28 published findings across many samples and settings: 15 of the 28 (54%) replicated, 75% of the 28 replication effect sizes were smaller than the originals, and the median Cohen's fell from 0.60 to 0.15 (Klein et al., 2018).
Sources cited in Section 37.1 3
- Cohen (1988) Statistical Power Analysis for the Behavioral Sciences
- Bagaïni et al. (2025) A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures
- Klein et al. (2018) Many Labs 2: Investigating Variation in Replicability Across Samples and Settings
37.2 Judgment and decision-making #
The view that preferences are constructed during elicitation, rather than retrieved from memory, has a long history in this field (Slovic, 1995). On it, preference reversals, in which two ways of asking the same question produce opposite answers, show that a single latent utility is a simplification; on the opposing view they are errors around a stable preference, which newer evidence partly supports. Oprea (2024) found the patterns associated with prospect theory in riskless "mirror" tasks and interprets them as an effect of complexity rather than of attitudes toward risk, a dispute conducted outside peer review that is unresolved (Oprea, 2025), and Gigerenzer (2018) argues that many reported biases come from treating random error as systematic, from small-sample statistics, or from counting reasonable inferences as mistakes. The effects most often cited against independent comparisons are fatigue, choice-induced change, context effects, and anchoring, in which an arbitrary number shifts later estimates (Tversky and Kahneman, 1974).
37.2.1 Self-control findings that faded #
The case for short sessions with breaks usually cites parole boards whose favorable rulings fell from about 65% to nearly zero between food breaks (Danziger et al., 2011), and ego depletion, the idea that self-control draws on a limited resource that an effortful first task uses up (Baumeister et al., 1998). Ego depletion faded in three large tests: , with a confidence interval that included zero, in a registered replication across 23 labs () (Hagger et al., 2016); a nonsignificant across 36 labs (), preregistered, with the data four times more probable under the null in a Bayesian analysis with an informed prior (Vohs et al., 2021); and a significant but small , or 0.16 after excluding participants who may have answered at random, across 12 labs () (Dang et al., 2021). Even if it exists, its size is around 0.1 and depends on the paradigm.
The parole-board effect did not survive scrutiny of its size: prisoners without legal representation tended to be heard last in each session (Weinshall-Margel and Shapard, 2011), and rational scheduling of cases alone can inflate the apparent effect (Glöckner, 2016). Decision fatigue, a decline in the quality of decisions over a run of them, met a large field test: in 231,076 triage calls handled by 174 nurses at Sweden's national healthcare advice line, four preregistered confirmatory tests all supported the null hypothesis, with one-sided Bayes factors above 22 (Andersson et al., 2025). A preregistered systematic review of 82 studies of healthcare professionals found a significant effect in 45% of quantitative tests, but notes that the construct is defined inconsistently and operationalized poorly (Maier et al., 2025). That people under high cognitive load chose chocolate cake over fruit salad more often (63% against 41%) (Shiv and Fedorikhin, 1999) is often described as replicated many times, but we found no registered or large-sample replication, and cognitive load did not change a large decoy effect in grocery choices (, within-subjects) (Wedell et al., 2022). The unconscious thought advantage, choosing better after distraction than after deliberation (Dijksterhuis et al., 2006), remains unsupported by a meta-analysis and large-scale replication (Nieuwenstein et al., 2015). Long sessions are not thereby harmless, but depletion and fatigue are no longer a sound basis for designing them.
37.2.2 Choice changes preference #
In the free-choice paradigm (Brehm, 1956), people rate items, choose between two they rated about equally, and rate everything again. The chosen item tends to rise and the rejected one to fall, a spreading of alternatives long read as the mind justifying its choice. Try it in Figure 37.1 before reading on.
Only part of your spread can be a change in preference. If a rating measures liking imperfectly and a choice follows true liking, then of two items rated alike, the one you choose is more likely to be the one your first rating underestimated, and a fresh rating will favor it even if your liking never changed; Chen and Risen (2010) proved this as a theorem and demonstrated it experimentally. Pooling four studies that address the critique, Izuma and Murayama (2013) found an effect that was statistically significant but small, . Figure 37.2 simulates the artifact.
At the classic design still shows a clear positive spread, as large as the control design's; set and only the classic spread rises, by about 0.4. Designs that separate the effect of choosing from what the choice reveals are called artifact-free. A meta-analysis of 43 studies with artifact-free free-choice designs () found , 95% confidence interval , with no evidence of publication bias, and concluded that choice causes real preference change, not merely reflects it (Enisman et al., 2021). Each choice raises the value of the chosen option and lowers that of the unchosen one, and choices between the same pair become less consistent as more trials intervene (Zylberberg et al., 2024); a sequential sampling model, in which evidence for each option accumulates noisily until it reaches a threshold (Section 39.3), reproduced the post-choice spreading of 457 participants (Lee and Pezzulo, 2026). Exposure matters too: options seen more often come to be liked more (Zajonc, 1968); across 268 curve estimates from 81 articles, liking follows an inverted U in the number of exposures, for visual but not auditory stimuli (Montoya et al., 2017); and what matters is relative exposure, how often an option is seen compared with the others (Mrkva and Van Boven, 2020). Figure 37.3 sets the effects of this section side by side.
Things to try:
- At the starting value, (36 labs of ego depletion), the two distributions overlap 98%, and a person drawn from the higher group outscores one from the lower group 52% of the time, barely better than a coin flip.
- Click the 43 artifact-free studies, : the overlap falls to 84%, and the higher group wins 61% of the time. The Many Labs 2 medians, 0.60 and 0.15, overlap 76% and 94%.
37.2.3 Context effects #
Luce's choice axiom (Section 16.4) says that the odds of choosing over do not depend on what else is offered. Three classic context effects violate it. Adding a decoy, an option worse than on every attribute but not worse than , raises the share choosing (the attraction effect) (Huber et al., 1982); an option gains when it becomes the middle one of three (compromise) (Simonson, 1989); and a new option takes its share mostly from the option it resembles (similarity) (Tversky, 1972). Try the first in Figure 37.4.
Every third option in round 2 was a decoy. In three situations it favored the option you passed over in round 1, so the attraction effect predicts a switch; in the other three it favored your choice, so a switch there measures the plain inconsistency of repeated choice (Section 37.4). With three answers each, chance alone produces either ordering often.
Since 2017 the evidence says that context effects hold at the population level under strong boundary conditions. Individuals usually do not show all three at once, and averaging across people can produce a pattern that no individual shows (Liew et al., 2016). A dominated option can also make the option that dominates it look worse, a repulsion effect, depending on how the stimuli are presented (Spektor et al., 2018). A review traced the appearance, disappearance, and reversal of context effects to the spatial arrangement of the options, how concrete the attributes are, and how long people deliberate (Spektor et al., 2021); the attraction effect appears mostly when every attribute is a number, and usually disappeared when people experienced the products or when even one attribute was presented perceptually (Frederick et al., 2014). A distractor effect attributed to divisive normalization, which divides each option's value by the total value on offer (Section 39.4), failed to replicate in two preregistered experiments (Gluth et al., 2020).
The most useful result for system design is about cues. In an incentivized, preregistered experiment with 909 participants and 40 trials each, every 10 percentage points of historical predictive power that a decoy cue had for the better option raised the probability of choosing the option the decoy favored by 1.29 percentage points; the corresponding figures were 1.58 for default options and 2.21 for an arbitrary rule (Forsgren et al., 2025b). People learn which cues point to good options and follow them. Anchoring is robust in groups, but a meta-analysis of more than 50,000 anchoring estimates found that the reliability of individual susceptibility is low in most tasks (Röseler et al., 2024), so a person's "anchoring score" cannot be measured well.
37.2.4 What this means for the observation model #
None of the following was tested in PBO; each is the book's inference.
- Put choice-induced updating in the likelihood. After beats , a model in the style of Zylberberg et al. (2024) updates and . Whether the free-choice magnitude transfers to design comparisons is untested; a simple check is to repeat a few early pairs late in the session and read a systematic flip toward the recently chosen option as drift (inference).
- The optimizer causes the drift it should guard against. The incumbent is compared most often, so it is the most exposed to revaluation and relative exposure; apparent convergence may be lock-in on an early winner (inference), one more entry for Section 19.6 and a candidate explanation for the unstable feedback of Section 32.3.
- Binary queries avoid the classic decoy, galleries do not. A gallery (Section 20.3), and options still in memory, reintroduce a context set in which, by the repulsion effect, the incumbent can distort how new candidates are perceived (inference).
- "New" is a learnable cue. An optimizer that keeps proposing better candidates teaches the person that the new option is usually better. Randomize where and in what order proposals appear, and occasionally compare two old options (inference).
- Set session length by measurement, not by depletion. Measure trends in lapse rates and response times rather than assume fatigue (inference).
- Do not fit per-user bias parameters. Individual susceptibility to anchoring and decoys is measured unreliably; a population-level bias term, or none, is safer (inference).
- Evaluate in use, not only in the session. We found no direct replication or new large-sample test since 2017 of the distinction between decision utility and experienced utility (Kahneman et al., 1997), so the outcome still needs a test in actual use, separate from the comparisons that produced it (inference), as Section 46.7 recommends.
Sources cited in Section 37.2 38
- Slovic (1995) The construction of preference
- Oprea (2024) Decisions under Risk Are Decisions under Complexity
- Oprea (2025) Initial Reply to Banki, Simonsohn, Walatka and Wu (2025)
- Gigerenzer (2018) The Bias Bias in Behavioral Economics
- Tversky and Kahneman (1974) Judgment under Uncertainty: Heuristics and Biases
- Danziger et al. (2011) Extraneous factors in judicial decisions
- Baumeister et al. (1998) Ego depletion: Is the active self a limited resource?
- Hagger et al. (2016) A Multilab Preregistered Replication of the Ego-Depletion Effect
- Vohs et al. (2021) A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect
- Dang et al. (2021) A Multilab Replication of the Ego Depletion Effect
- Weinshall-Margel and Shapard (2011) Overlooked factors in the analysis of parole decisions
- Glöckner (2016) The irrational hungry judge effect revisited: Simulations reveal that the magnitude of the effect is overestimated
- Andersson et al. (2025) No evidence for decision fatigue using large-scale field data from healthcare
- Maier et al. (2025) Systematic review of the effects of decision fatigue in healthcare professionals on medical decision-making
- Shiv and Fedorikhin (1999) Heart and Mind in Conflict: the Interplay of Affect and Cognition in Consumer Decision Making
- Wedell et al. (2022) Context effects on choice under cognitive load
- Dijksterhuis et al. (2006) On Making the Right Choice: The Deliberation-Without-Attention Effect
- Nieuwenstein et al. (2015) On making the right choice: A meta-analysis and large-scale replication attempt of the unconscious thought advantage
- Brehm (1956) Postdecision changes in the desirability of alternatives
- Chen and Risen (2010) How choice affects and reflects preferences: Revisiting the free-choice paradigm
- Izuma and Murayama (2013) Choice-Induced Preference Change in the Free-Choice Paradigm: A Critical Methodological Review
- Enisman et al. (2021) Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm
- Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
- Lee and Pezzulo (2026) Choice-induced preference change under a sequential sampling model framework
- Zajonc (1968) Attitudinal effects of mere exposure
- Montoya et al. (2017) A re-examination of the mere exposure effect: The influence of repeated exposure on recognition, familiarity, and liking
- Mrkva and Van Boven (2020) Salience theory of mere exposure: Relative exposure increases liking, extremity, and emotional intensity
- Huber et al. (1982) Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis
- Simonson (1989) Choice Based on Reasons: The Case of Attraction and Compromise Effects
- Tversky (1972) Elimination by aspects: A theory of choice
- Liew et al. (2016) The appropriacy of averaging in the study of context effects
- Spektor et al. (2018) When the Good Looks Bad: An Experimental Exploration of the Repulsion Effect
- Spektor et al. (2021) The elusiveness of context effects in decision making
- Frederick et al. (2014) The Limits of Attraction
- Gluth et al. (2020) Value-based attention but not divisive normalization influences decisions with multiple alternatives
- Forsgren et al. (2025b) Probabilistic functionalism as a limiting condition for robustness
- Röseler et al. (2024) Measurements of Susceptibility to Anchoring are Unreliable: Meta-Analytic Evidence From More Than 50,000 Anchored Estimates
- Kahneman et al. (1997) Back to Bentham? Explorations of Experienced Utility
37.3 Psychophysics and attention #
Psychophysics supplied the first models of comparison (Section 16.2), and two common claims about PBO come from it. By Weber's law, the smallest noticeable difference grows with the magnitude being judged (Fechner, 1860), and judged magnitude follows a power law of physical magnitude (Stevens, 1957), so as the optimizer converges the person may no longer tell the candidates apart. And by decision by sampling, the subjective value of an amount is its rank among amounts sampled from memory and from the current context (Stewart et al., 2006), so an optimizer that concentrates its queries on good options would lead the person to undervalue them. A further claim cites the limit on absolute judgment, 2.6 bits on average along a single dimension (Miller, 1956), as strong evidence that comparisons beat ratings.
37.3.1 Relative value: sampling, range, and normalization #
For described attributes such as amounts and probabilities, decision by sampling weakened. A multi-lab, quasi-adversarial replication (proponents and skeptics agree on the design in advance) reproduced the effect of the distribution of values presented on the shapes of the utility and probability-weighting functions, but also observed it where decision by sampling predicts no effect, and showed by simulation that a misspecified choice model can produce it (Alempaki et al., 2019); a preregistered falsification test then found strong evidence against the prediction that manipulating the ranks of the amounts changes choices (Forsgren et al., 2025a).
For values learned from experienced outcomes, range adaptation is well supported. Human reinforcement learning is best fitted by models that center values on a reference point and adapt them to the range of outcomes (Bavard et al., 2018); the adaptation causes systematic errors when values are extrapolated to new contexts, larger when the task is easier (Bavard et al., 2021); and a task built to tell the two candidate rules apart falsified divisive normalization and supported range normalization, in which a value is rescaled by the difference between the best and worst values in its context (Bavard and Palminteri, 2023). Figure 37.5 computes what that would do to a session. A range-normalized person experiences, in session ,
where and are the worst and best utilities shown and between 0 and 1 sets how strongly they adapt; ratings are , and comparisons follow the probit model on the experienced scale.
Things to try:
- With and session B near the best, the probability of preferring is much closer to 1 in session B, so a model fitted there would learn a larger amplitude. At the sessions agree.
- Move a design outside the shaded region: session B extrapolates, the regime in which Bavard et al. (2021) found systematic errors.
37.3.2 Sequential effects #
A judgment can be pulled toward the previous one (assimilation) or pushed away from it (contrast); the direction depends on the setting, and neither is a stable personal trait. In visual perception, the positive pull decays over time, requires the features to be similar, depends on spatial location, and is modulated by attention (Manassi et al., 2023). Preference ratings of photographs and faces assimilate toward the previous stimulus even after response bias is controlled (Chang et al., 2017), while in 2.2 million Yelp ratings and 4.2 million Amazon ratings the same reviewer's ratings show contrast relative to their previous ratings (Vinson et al., 2019). Assimilation in ratings of facial attractiveness is not stable in the same person across two sequences (Kramer and Cartledge, 2026), and in repeated aesthetic ratings both assimilation and contrast grow with the similarity of the stimuli (Pombo et al., 2023).
37.3.3 Attention #
That attention affects choice holds in direction, but whether it raises value is disputed. Across six eye-tracking data sets, the sum of the options' values affected response times and, in most data sets, the relation between gaze and choice, consistent with gaze multiplying value (Smith and Krajbich, 2019). A review concluded that there is not enough evidence that attention itself raises the perceived value of an option (Mormann and Russo, 2021). When only one option is visible at a time, the attentional choice bias roughly doubles () (Eum et al., 2023). Section 39.3 reports a meta-analysis of the causal effect of attention.
37.3.4 What this means for the observation model #
- The learned utility is on a session-relative scale (Figure 37.5). Comparing posteriors across sessions or users, or warm-starting from an old posterior, needs explicit recalibration, and extrapolation beyond the explored range is where errors are most likely (inference).
- For perceptual options, expect range and recency, not curvature. Range adaptation and sequential dependence change the scale and recent-trial bias rather than the shape of the utility (inference).
- Present simultaneously and balance positions. Show both candidates side by side, balance left and right, and do not make the system's proposal more salient; where options can only be experienced in sequence (audio, animation), counterbalance order across repeated pairs (inference). Section 20.6 makes the general case, and the photo-enhancement case study (Chapter 25) is the kind of perceptual task where these effects should be expected.
- Model recency at the population level, and estimate its direction. A per-user sequential-bias parameter is unstable within a person, and the sign, assimilation or contrast, should be estimated rather than assumed (inference).
- Discrimination near the optimum depends on what is judged. For value judgments the Weber-based claim fails: decisions between two high-value options tend to be faster and more accurate (Shevlin et al., 2022), and discrimination tracks how densely values occur rather than how large they are (Polanía et al., 2019) (Section 39.4). For the physical parameters of a device, such as the stiffness of a prosthetic ankle, a discrimination floor does exist (Shepherd et al., 2018; Maberry and Martin, 2026) (Section 39.7, Chapter 24).
Two common claims need correcting. Luce's axiom and Thurstone's Case V are not equivalent "under a logistic distribution": Case V assumes normal errors (the probit link, Section 16.3), and Yellott (1977) shows that the choice axiom follows from independent double exponential (Gumbel) errors (the logistic link, Section 16.5); for pairs other distributions do the same, while for choices among three options the double exponential is the only one. And we found no test of Miller's channel-capacity argument with design stimuli.
Sources cited in Section 37.3 22
- Fechner (1860) Elemente der Psychophysik
- Stevens (1957) On the Psychophysical Law
- Stewart et al. (2006) Decision by sampling
- Miller (1956) The magical number seven, plus or minus two: Some limits on our capacity for processing information
- Alempaki et al. (2019) Reexamining How Utility and Weighting Functions Get Their Shapes: A Quasi-Adversarial Collaboration Providing a New Interpretation
- Forsgren et al. (2025a) A Preregistered Falsification Test of the Decision by Sampling Model and Rank-Order Effect
- Bavard et al. (2018) Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences
- Bavard et al. (2021) Two sides of the same coin: Beneficial and detrimental consequences of range adaptation in human reinforcement learning
- Bavard and Palminteri (2023) The functional form of value normalization in human reinforcement learning
- Manassi et al. (2023) Serial dependence in visual perception: A meta-analysis and review
- Chang et al. (2017) Sequential effects in preference decision: Prior preference assimilates current preference
- Vinson et al. (2019) Decision contamination in the wild: Sequential dependencies in online review ratings
- Kramer and Cartledge (2026) Sequential effects in facial attractiveness judgements: No evidence of stable individual differences
- Pombo et al. (2023) The intrinsic variance of beauty judgment
- Smith and Krajbich (2019) Gaze Amplifies Value in Decision Making
- Mormann and Russo (2021) Does Attention Increase the Value of Choice Alternatives?
- Eum et al. (2023) Peripheral Visual Information Halves Attentional Choice Biases
- Shevlin et al. (2022) High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity
- Polanía et al. (2019) Efficient coding of subjective value
- Shepherd et al. (2018) Amputee perception of prosthetic ankle stiffness during locomotion
- Maberry and Martin (2026) Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton
- Yellott (1977) The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution
37.4 Mathematical psychology #
Mathematical psychology tests the formal models from which the probit and logistic links come. Random utility theory (Section 16.5) has long allowed the noise to be read as variation within a person, variation between people, or genuine instability, readings that are formally interchangeable (McFadden, 1981). A preference is transitive if preferring to and to implies preferring to ; since Tversky (1969) reported violations, many have been claimed, and Regenwetter et al. (2011) concluded that most of them can be explained by people mixing among transitive preferences. With noisy choices, every model in which the probability of a choice depends on a utility difference, the probit model of Equation (27.1) included, satisfies strong stochastic transitivity: if beats and beats at least half the time, beats at least as often as in either pair, because is at least as large as either difference. A violation of even the weak form, beating less than half the time, is evidence against the whole family.
Quantum cognition models judgments with the probability rules of quantum mechanics, where asking question 1 and then question 2 can give different answers than the reverse order. A parameter-free prediction about such order effects, the QQ equality, was supported in 70 national surveys and two laboratory experiments (Wang et al., 2014). From this comes the argument that each query changes the state of a person's preference, that repeated queries "may never converge" to a true preference, and that PBO should use a surrogate built on quantum probability.
37.4.1 Random utility, tested directly #
Direct tests support the random utility hypothesis. Using Falmagne's inequalities, which characterize exactly which choice probabilities a random utility model can produce, McCausland et al. (2020) tested 141 participants, each of whom chose six times from every subset of at least two of five lotteries. Most participants satisfied random utility; only 4 showed strong evidence of violating it.
The way the noise is specified, however, changes the inference. Combining noise in the preference parameters with noise in the response process can make an expected-value maximizer look risk averse or risk seeking, so modal choices cannot simply be used to infer latent preferences (Bhatia and Loomes, 2017). When noise of one scale is added to the utilities of every pair, the probability of choosing the riskier option need not fall as risk aversion grows, which makes risk aversion hard to identify; random parameter (random preference) models do not have this problem (Apesteguia and Ballester, 2018). For example (computed for this book, not from their data), with a lottery paying 80 or 10 against a sure 40, power utility, and probit noise of scale 1, a person with risk-aversion parameter picks the lottery with probability 0.29 and a more risk-averse one with with probability 0.49: a large compresses all utility differences, and a fixed noise scale turns them into coin flips. A 2024 preprint qualifies the criticism: in a representative Danish sample (253 people), the standard expected-utility specification squeezed estimates of risk aversion into a narrow range when everyone shared one noise scale, and gave estimates in line with other models when each person had their own (Keffert and Schweizer, 2024).
37.4.2 Transitivity and cycles #
Response times can separate noisy cycles from genuine ones: choices between options of similar value are both slower and more error-prone, so a cycle made only of fast choices is hard to explain as noise. In two reanalyzed data sets, a working paper found that 54.38% and 90.25% of the violations of weak stochastic transitivity involved only fast choices, and that 19.24% and 39.58% of them qualified under a stricter criterion, in which choice frequencies and response times reveal the preference on each pair of the cycle; for the rest, neither noise nor intransitive preference can be ruled out (Alós-Ferrer et al., 2023). Violations shrink but do not disappear once noise is separated from preference, and the most frequent arise from chains of small trade-offs. Lotteries designed after the Steinhaus-Trybula paradox still produced cycles as the most common pattern after random-but-transitive explanations were accounted for (Butler and Pogrebna, 2018).
37.4.3 Order effects and quantum models #
Quantum cognition holds up in part, but its claim to uniqueness is contested. Classical repeat-choice models also yield the QQ equality as a parameter-free prediction and account for the survey data comparably (Kellen et al., 2018). In an experiment with 325 participants and 12 sets of issues, inserting an incompatible question between two presentations of the same question made the answer to the repeated question uncertain, as quantum probability predicts (Busemeyer and Wang, 2017), though a 2023 preprint argues that this experiment is logically inconsistent with results on the replicability of answers (Ozawa and Khrennikov, 2023). Kvam et al. (2021) found that the strength of a preference oscillates over time and that eliciting a choice earlier changes the later oscillation.
37.4.4 What this means for the observation model #
- The standard likelihood is a choice, not a neutral default. A fixed-scale probit or logistic link is a homoscedastic random utility model and can distort the inferred strength of a preference. Alternatives are a random preference likelihood, which draws each comparison from one whole sample of the Gaussian process posterior, and a heteroscedastic link whose noise scale varies with the difficulty of the pair (inference; shaped by the Gaussian process framework). At the least, report the reversal rate on a few repeated pairs next to any fitted noise scale (inference).
- Local cycles may need modeling, general intransitivity rarely does. Chains of small trade-offs describe exactly the local queries around the incumbent late in a session; a mixture of rankings or a random preference model can absorb them without giving up a latent utility (inference).
- Retest immediately or much later, not in between. By the interleaved-question and oscillation results, a repeat separated by other queries is contaminated. Retest right away to measure response noise, or much later to measure drift (inference).
- Quantum surrogates are not a default. They are untested in any optimization setting, and the order effects behind them have classical explanations; a covariate for query order is simpler, and "converging to a point that depends on the session" is closer to the evidence than "may never converge" (inference).
- Better simulated users. Differentiable decision theories fitted to the largest experiment on risky choice to date, more accurate than the classic ones (Peterson et al., 2021), and a language model fine-tuned on 10 million choices from 160 experiments, which predicted held-out participants better than existing cognitive models (Binz et al., 2025), could serve as simulated users (Section 31.7), although how well they fit design comparisons is unknown (inference).
Sources cited in Section 37.4 16
- McFadden (1981) Econometric Models of Probabilistic Choice
- Tversky (1969) Intransitivity of preferences
- Regenwetter et al. (2011) Transitivity of preferences
- Wang et al. (2014) Context effects produced by question orders reveal quantum nature of human judgments
- McCausland et al. (2020) Testing the Random Utility Hypothesis Directly
- Bhatia and Loomes (2017) Noisy preferences in risky choice: A cautionary note
- Apesteguia and Ballester (2018) Monotone Stochastic Choice Models: The Case of Risk and Time Preferences
- Keffert and Schweizer (2024) Stochastic Monotonicity and Random Utility Models: The Good and The Ugly
- Alós-Ferrer et al. (2023) Identifying Nontransitive Preferences
- Butler and Pogrebna (2018) Predictably intransitive preferences
- Kellen et al. (2018) Classic-probability accounts of mirrored (quantum-like) order effects in human judgments
- Busemeyer and Wang (2017) Is there a problem with quantum models of psychological measurements?
- Ozawa and Khrennikov (2023) The logical inconsistency of the model and experiment presented in the paper of Busemeyer and Wang “Is there a problem with quantum models of psychological measurements?”
- Kvam et al. (2021) Temporal oscillations in preference strength provide evidence for an open system model of constructed preference
- Peterson et al. (2021) Using large-scale experiments and machine learning to discover theories of human decision-making
- Binz et al. (2025) A foundation model to predict and capture human cognition
37.5 Ecological rationality and heuristics #
The research program on ecological rationality studies simple decision rules, heuristics, and the environments in which they work. Take-the-best, which decides on the most valid cue that discriminates between the options, predicted out of sample no worse than multiple regression across 20 data sets, and the bias-variance trade-off explains why less information can lead to better predictions (Gigerenzer and Brighton, 2009). From this comes the argument that simpler surrogates may do better when data are sparse, and that a smooth compensatory utility, in which a loss on one attribute can be made up by a gain on another, misdescribes a person who uses a lexicographic rule.
37.5.1 When people use which strategy #
Since 2017 the question has moved from whether heuristics work to when each is used. Under time pressure take-the-best is slower and less accurate than tallying cues, and the result reverses when the stimulus format makes its search cheaper (Bobadilla-Suarez and Love, 2018). Bounded meta-learned inference predicts that people decide by a single cue when they know the ranking of the attributes' importance, weigh attributes equally when they know only the direction of each, and combine weighted attributes when they know neither; three experiments with pairwise comparisons of options described by continuous features supported these predictions (Binz et al., 2022). On very large risky-choice data, more expressive but still interpretable models beat the classic theories (Peterson et al., 2021), so "less is more" depends on how much data there is, as the bias-variance argument implies.
37.5.2 What this means for the observation model #
The choice rule can change while the utility does not. At the start of a session, a person usually does not know which design attribute matters, a state that the meta-learning theory maps to weighted multi-attribute comparison. As the session teaches them which attribute decides, the same theory predicts a shift toward single-reason, lexicographic choice, so a smooth compensatory utility would fit the early session better than the late one. This is non-stationarity of the choice rule, distinct from drift in the utility itself, and the two should be modeled separately (inference); Section 46.5 takes up both. The interface matters too: listing parameter values side by side makes single-reason strategies cheaper, while a holistic rendering favors integrated judgment (inference).
A lexicographic person within a smooth model. A continuous latent utility can only approximate a lexicographic ordering. A repair inside the framework is to allow a much shorter lengthscale (Section 9.2) along the attribute the person has locked onto, or a function close to a step (inference; this repair comes from the Gaussian process framework, not from the problem itself). In a six-parameter design problem, a person who has decided that one parameter matters most would show a short lengthscale along it and long ones along the other five, a shape that one lengthscale per input can represent and a single shared lengthscale cannot (inference).
Simpler surrogates and bias terms need evidence. Whether a low-dimensional or additive utility beats a full Gaussian process on human pairwise data from real design sessions is untested. The "bias bias" critique and the complexity results counsel caution before adding directional bias terms: each needs evidence that it is systematic in the task at hand rather than random error (inference).
Sources cited in Section 37.5 4
- Gigerenzer and Brighton (2009) Homo Heuristicus: Why Biased Minds Make Better Inferences
- Bobadilla-Suarez and Love (2018) Fast or frugal, but not both: Decision heuristics under time pressure
- Binz et al. (2022) Heuristics from bounded meta-learned inference
- Peterson et al. (2021) Using large-scale experiments and machine learning to discover theories of human decision-making
37.6 Settled, contested, missing #
Settled. Ego depletion, if it exists, is around (Vohs et al., 2021; Dang et al., 2021); a preregistered field test of 231,076 calls found no decision fatigue (Andersson et al., 2025). Choice changes preference: across 43 artifact-free studies (Enisman et al., 2021). Context effects exist at the population level but appear, vanish, and reverse with presentation, and individual susceptibility to anchoring cannot be measured reliably. Most people's repeated choices satisfy random utility (McCausland et al., 2020). For values learned from experience, range normalization fits better than divisive normalization (Bavard and Palminteri, 2023).
Contested. Whether the remaining reports of decision fatigue in healthcare reflect a real effect or a poorly defined construct. Whether patterns attributed to risk preference are effects of complexity (Oprea, 2024). How much observed intransitivity is noise: a working paper finds that part of it survives once noise is separated from preference (Alós-Ferrer et al., 2023). Whether quantum models are needed for order effects. Which heuristic people use in design tasks.
Missing. Any test of whether the size of choice-induced change carries over to comparisons of designs. A PBO study that models choice-induced updating, relative exposure, or the "new is better" cue, or that retests early pairs late in a session. A comparison of homoscedastic, heteroscedastic, and random preference likelihoods on human design comparisons. Taken together, the evidence favors an observation model with a stable component, structured evaluation noise, and query-induced drift, which no preferential optimization method yet implements (inference).
Sources cited in Section 37.6 8
- Vohs et al. (2021) A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect
- Dang et al. (2021) A Multilab Replication of the Ego Depletion Effect
- Andersson et al. (2025) No evidence for decision fatigue using large-scale field data from healthcare
- Enisman et al. (2021) Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm
- McCausland et al. (2020) Testing the Random Utility Hypothesis Directly
- Bavard and Palminteri (2023) The functional form of value normalization in human reinforcement learning
- Oprea (2024) Decisions under Risk Are Decisions under Complexity
- Alós-Ferrer et al. (2023) Identifying Nontransitive Preferences
Further reading #
- Chen and Risen (2010) is a short lesson in how a measurement design can manufacture an effect; read it with the artifact-free meta-analysis of Enisman et al. (2021).
- Spektor et al. (2021) review when context effects appear, vanish, and reverse, the best single source for anyone designing a multi-option interface.
- Klein et al. (2018) shows what happened to 28 classic findings run across many labs, a calibration for every effect size in this part.
- McCausland et al. (2020) and Apesteguia and Ballester (2018) together explain how to test a random utility model and why its noise specification matters.
References
- (2019). Reexamining How Utility and Weighting Functions Get Their Shapes: A Quasi-Adversarial Collaboration Providing a New Interpretation. Management Science. Cited in §37.3
- (2023). Identifying Nontransitive Preferences. University of Zurich. working paper Cited in §37.4 §37.6
- (2025). No evidence for decision fatigue using large-scale field data from healthcare. Communications Psychology. Cited in §37.2 §37.6
- (2018). Monotone Stochastic Choice Models: The Case of Risk and Time Preferences. Journal of Political Economy. Cited in §37.4
- (2025). A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures. Nature Human Behaviour. doi:10.1038/s41562-024-02085-2. Cited in §37.1
- (1998). Ego depletion: Is the active self a limited resource? Journal of Personality and Social Psychology. Cited in §37.2
- (2023). The functional form of value normalization in human reinforcement learning. eLife. Cited in §37.3 §37.6
- (2018). Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences. Nature Communications. Cited in §37.3
- (2021). Two sides of the same coin: Beneficial and detrimental consequences of range adaptation in human reinforcement learning. Science Advances. Cited in §37.3
- (2017). Noisy preferences in risky choice: A cautionary note. Psychological Review. Cited in §37.4
- (2022). Heuristics from bounded meta-learned inference. Psychological Review. Cited in §37.5
- (2025). A foundation model to predict and capture human cognition. Nature. Cited in §37.4
- (2018). Fast or frugal, but not both: Decision heuristics under time pressure. Journal of Experimental Psychology: Learning, Memory, and Cognition. Cited in §37.5
- (1956). Postdecision changes in the desirability of alternatives. The Journal of Abnormal and Social Psychology. Cited in §37.2
- (2017). Is there a problem with quantum models of psychological measurements? PLOS ONE. Cited in §37.4
- (2018). Predictably intransitive preferences. Judgment and Decision Making. Cited in §37.4
- (2017). Sequential effects in preference decision: Prior preference assimilates current preference. PLOS ONE. Cited in §37.3
- (2010). How choice affects and reflects preferences: Revisiting the free-choice paradigm. Journal of Personality and Social Psychology. Cited in §37.2
- (1988). Statistical Power Analysis for the Behavioral Sciences. Lawrence Erlbaum Associates. Cited in §37.1
- (2021). A Multilab Replication of the Ego Depletion Effect. Social Psychological and Personality Science. Cited in §37.2 §37.6
- (2011). Extraneous factors in judicial decisions. Proceedings of the National Academy of Sciences. Cited in §37.2
- (2006). On Making the Right Choice: The Deliberation-Without-Attention Effect. Science. Cited in §37.2
- (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Cited in §37.2 §37.6
- (2023). Peripheral Visual Information Halves Attentional Choice Biases. Psychological Science. Cited in §37.3
- (1860). Elemente der Psychophysik. Breitkopf und Härtel. Cited in §37.3
- (2025a). A Preregistered Falsification Test of the Decision by Sampling Model and Rank-Order Effect. Management Science. Cited in §37.3
- (2025b). Probabilistic functionalism as a limiting condition for robustness. Scientific Reports. Cited in §37.2
- (2014). The Limits of Attraction. Journal of Marketing Research. Cited in §37.2
- (2018). The Bias Bias in Behavioral Economics. Review of Behavioral Economics. Cited in §37.2
- (2009). Homo Heuristicus: Why Biased Minds Make Better Inferences. Topics in Cognitive Science. Cited in §37.5
- (2016). The irrational hungry judge effect revisited: Simulations reveal that the magnitude of the effect is overestimated. Judgment and Decision Making. Cited in §37.2
- (2020). Value-based attention but not divisive normalization influences decisions with multiple alternatives. Nature Human Behaviour. Cited in §37.2
- (2016). A Multilab Preregistered Replication of the Ego-Depletion Effect. Perspectives on Psychological Science. Cited in §37.2
- (1982). Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis. Journal of Consumer Research. Cited in §37.2
- (2013). Choice-Induced Preference Change in the Free-Choice Paradigm: A Critical Methodological Review. Frontiers in Psychology. Cited in §37.2
- (1997). Back to Bentham? Explorations of Experienced Utility. The Quarterly Journal of Economics. Cited in §37.2
- (2024). Stochastic Monotonicity and Random Utility Models: The Good and The Ugly. arXiv preprint 2409.00704. preprint Cited in §37.4
- (2018). Classic-probability accounts of mirrored (quantum-like) order effects in human judgments. Decision. Cited in §37.4
- (2018). Many Labs 2: Investigating Variation in Replicability Across Samples and Settings. Advances in Methods and Practices in Psychological Science. Cited in §37.1
- (2026). Sequential effects in facial attractiveness judgements: No evidence of stable individual differences. Perception. Cited in §37.3
- (2021). Temporal oscillations in preference strength provide evidence for an open system model of constructed preference. Scientific Reports. Cited in §37.4
- (2026). Choice-induced preference change under a sequential sampling model framework. Scientific Reports. doi:10.1038/s41598-026-44610-5. Cited in §37.2
- (2016). The appropriacy of averaging in the study of context effects. Psychonomic Bulletin & Review. Cited in §37.2
- (2026). Just Noticeable Difference of Impedance Parameters While Walking in an Ankle Exoskeleton. IEEE Transactions on Neural Systems and Rehabilitation Engineering. Cited in §37.3
- (2025). Systematic review of the effects of decision fatigue in healthcare professionals on medical decision-making. Health Psychology Review. Cited in §37.2
- (2023). Serial dependence in visual perception: A meta-analysis and review. Journal of Vision. Cited in §37.3
- (2020). Testing the Random Utility Hypothesis Directly. The Economic Journal. doi:10.1093/ej/uez039. Cited in §37.4 §37.6
- (1981). Econometric Models of Probabilistic Choice. Structural Analysis of Discrete Data with Econometric Applications. Cited in §37.4
- (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review. Cited in §37.3
- (2017). A re-examination of the mere exposure effect: The influence of repeated exposure on recognition, familiarity, and liking. Psychological Bulletin. Cited in §37.2
- (2021). Does Attention Increase the Value of Choice Alternatives? Trends in Cognitive Sciences. Cited in §37.3
- (2020). Salience theory of mere exposure: Relative exposure increases liking, extremity, and emotional intensity. Journal of Personality and Social Psychology. Cited in §37.2
- (2015). On making the right choice: A meta-analysis and large-scale replication attempt of the unconscious thought advantage. Judgment and Decision Making. Cited in §37.2
- (2024). Decisions under Risk Are Decisions under Complexity. American Economic Review. Cited in §37.2 §37.6
- (2025). Initial Reply to Banki, Simonsohn, Walatka and Wu (2025). Data Colada. non-peer-reviewed Cited in §37.2
- (2023). The logical inconsistency of the model and experiment presented in the paper of Busemeyer and Wang “Is there a problem with quantum models of psychological measurements?”. PsyArXiv. preprint Cited in §37.4
- (2021). Using large-scale experiments and machine learning to discover theories of human decision-making. Science. Cited in §37.4 §37.5
- (2019). Efficient coding of subjective value. Nature Neuroscience. Cited in §37.3
- (2023). The intrinsic variance of beauty judgment. Attention, Perception, & Psychophysics. Cited in §37.3
- (2011). Transitivity of preferences. Psychological Review. Cited in §37.4
- (2024). Measurements of Susceptibility to Anchoring are Unreliable: Meta-Analytic Evidence From More Than 50,000 Anchored Estimates. Meta-Psychology. Cited in §37.2
- (2018). Amputee perception of prosthetic ankle stiffness during locomotion. Journal of NeuroEngineering and Rehabilitation. Cited in §37.3
- (2022). High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity. Proceedings of the National Academy of Sciences. Cited in §37.3
- (1999). Heart and Mind in Conflict: the Interplay of Affect and Cognition in Consumer Decision Making. Journal of Consumer Research. Cited in §37.2
- (1989). Choice Based on Reasons: The Case of Attraction and Compromise Effects. Journal of Consumer Research. Cited in §37.2
- (1995). The construction of preference. American Psychologist. Cited in §37.2
- (2019). Gaze Amplifies Value in Decision Making. Psychological Science. Cited in §37.3
- (2018). When the Good Looks Bad: An Experimental Exploration of the Repulsion Effect. Psychological Science. Cited in §37.2
- (2021). The elusiveness of context effects in decision making. Trends in Cognitive Sciences. Cited in §37.2
- (1957). On the Psychophysical Law. Psychological Review. Cited in §37.3
- (2006). Decision by sampling. Cognitive Psychology. Cited in §37.3
- (1969). Intransitivity of preferences. Psychological Review. Cited in §37.4
- (1972). Elimination by aspects: A theory of choice. Psychological Review. Cited in §37.2
- (1974). Judgment under Uncertainty: Heuristics and Biases. Science. Cited in §37.2
- (2019). Decision contamination in the wild: Sequential dependencies in online review ratings. Behavior Research Methods. Cited in §37.3
- (2021). A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect. Psychological Science. Cited in §37.2 §37.6
- (2014). Context effects produced by question orders reveal quantum nature of human judgments. Proceedings of the National Academy of Sciences. Cited in §37.4
- (2022). Context effects on choice under cognitive load. Psychonomic Bulletin & Review. Cited in §37.2
- (2011). Overlooked factors in the analysis of parole decisions. Proceedings of the National Academy of Sciences. Cited in §37.2
- (1977). The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution. Journal of Mathematical Psychology. Cited in §37.3
- (1968). Attitudinal effects of mere exposure. Journal of Personality and Social Psychology. Cited in §37.2
- (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Cited in §37.2