Why Ask for Comparisons
Every method in Part III assumed that evaluating the objective returns a number: a validation accuracy, a yield, a walking speed. When the objective is a person's judgment, the obvious number is a rating, "how good is this, from 1 to 10?", and the obvious plan is to feed ratings to the Gaussian process of Chapter 8 as if they were measurements. Section 1.4 warned that this plan runs into trouble, because people are poor at producing such numbers and much better at saying which of two things they prefer.
This chapter makes that warning precise. It starts with an experiment you run on yourself, then turns what you observe into a model: a comparison is a noisy measurement of the difference between two hidden utilities, and a choice model, the rule that maps that difference to the probability of each answer, is the likelihood that every later chapter in this part uses.
Chapter 17 then deals with the mathematical cost of using these likelihoods, Chapter 18 attaches them to a Gaussian process, and Chapter 19 uses the result to optimize.
16.1 People as the objective #
Start with the figure below. It asks you to judge gray patches in two ways. In the rating task one patch appears on its own, on a dark or a light surround, and you give it a number from 1 (darkest) to 9 (lightest). In the comparison task two patches appear side by side and you say which is lighter. Do a dozen or more trials of each, honestly and without going back, then read the panel at the bottom.
The bar chart asks both tasks one question, so they can be compared directly. A rating does not answer it on its own: two ratings, given at different moments to different patches, have to be set side by side after the fact. A comparison answers it in a single trial. The simulated observer shows the pattern the rest of this section explains. Its ratings put patches one level apart in the right order only about seven times in ten, while its comparisons almost never err. Its numbers are illustrative, not measured: each rating adds to the true level a shift of 0.6 levels (up on the dark surround, down on the light one), a slow drift in how the scale is used, and noise with a standard deviation of 0.8 levels, and is then rounded to the scale; each comparison sees only noise of 0.5 levels on the difference, because the surround and the drift are shared by both patches. Your own numbers will be noisier with a dozen trials, but the gap is usually easy to see, because the two tasks ask different things of you.
16.1.1 Why ratings are hard #
The first reason is that people can tell apart far fewer levels in isolation than side by side. Miller (1956) collected experiments in which listeners and viewers had to identify a single stimulus by giving it a number, as in the rating task above. For the pitch of a tone, the information that got through leveled off at about 2.5 bits, the equivalent of about six levels that a listener never confuses; across all the one-dimensional attributes he reviewed the mean was 2.6 bits, with a standard deviation of only 0.6 bit. The nine levels in Figure 16.1 are more than the six or so that Miller found typical, yet any two adjacent levels are easy to tell apart when shown together. The first remedy Miller listed for this limit was to make relative rather than absolute judgments.
The second reason is that a number depends on its context. The patch is the same gray whatever surrounds it, yet a gray on a dark surround looks lighter than the same gray on a light one, the simultaneous contrast that Thurstone already noted for gray values (Thurstone, 1927); the scatter in the figure shows whether your ratings shifted with the surround. Ratings of real things show the same dependence on what came just before. Preference ratings of photographs and faces are pulled toward the rating of the previous item, and the pull survives controls for response bias (Chang et al., 2017). In 2.2 million Yelp ratings and 4.2 million Amazon ratings, a reviewer's rating was pushed away from their previous ratings, a contrast effect (Vinson et al., 2019). Attractiveness ratings are pulled toward the previous face on average, but how strongly a given person is pulled is not stable from one sequence to the next (Kramer and Cartledge, 2026). The direction depends on the setting; the dependence itself is reliable.
The third reason is that people use the scale unevenly. More than a century ago Hollingworth (1910) described the central tendency of judgment: estimates of magnitude drift toward the middle of the range of stimuli a person has been shown, so small values are overestimated and large ones underestimated. A rating scale is likely to compress toward its middle in the same way (inference).
These effects appear in exactly the settings this book is about. In a three-month deployment of a preference-guided optimizer for processing 3D meshes, in which two professional artists rated candidate meshes, the optimization lacked mechanisms to deal with "inconsistent and contradictory human judgments", and what the system showed influenced later answers through heuristic biases and loss aversion; interviews pointed to anchoring on meshes seen earlier and to judgments losing precision after a run of increasingly good results (Ou et al., 2022). Koyama and Igarashi (2018) argue that an absolute rating requires familiarity with the whole design space, which a person meeting the space for the first time does not have, while a comparison between two options can be answered at once. Chapter 25 works through one such problem end to end, adjusting a photograph by choosing between versions.
16.1.2 Why comparisons help #
A comparison does not make these effects disappear. It arranges for many of them to cancel. If your mood, the surround, or your sense of where "5" sits on the scale shifts both options by the same amount, the difference between them is untouched, and a comparison reports only the sign of that difference. Thurstone made this argument in 1927: looking at two handwriting specimens "in a mood slightly more generous and tolerant than ordinarily" raises the impression of both, and to that extent the two impressions vary together (Thurstone, 1927). Section 16.3 shows how this shared variation drops out of the comparison.
A rating mixes the person's preference with everything that sets their scale at that moment. A comparison reports the sign of a difference, and a difference cancels whatever shifts both options equally.
Direct evidence that people compare more reliably than they rate comes from several fields. In information retrieval, where assessors judge which of two documents better answers a query, preference judgments are made faster and more consistently than graded judgments (Clarke et al., 2021). A 2014 preprint reports experiments on Amazon Mechanical Turk across a variety of tasks in which pairwise comparisons had lower noise per answer and were typically faster to collect than numerical scores, though each answer carried less information (Shah et al., 2014). In a study of how people report preferences in markets, participants found it harder to report cardinal information (how much they prefer something) than ordinal information (which they prefer) (Budish and Kessler, 2022).
The advantage is not universal, and it is worth knowing where it fails. When 162 people repeated a health valuation task two weeks apart, 76.4% of their discrete choices were the same (kappa 0.528, a measure of agreement corrected for chance, where 1 is perfect), while the numerical time trade-off method reached an intraclass correlation of 0.958 (the share of the variance that comes from differences between people, not between occasions) even though only 59.3% of its values were identical (Xie et al., 2022). For risk preference, a meta-analysis of test-retest correlations found that self-reported propensity to take risks was more stable over time than behavioral measures such as lottery choices, with estimated reliabilities of 0.61 and 0.25 (Bagaïni et al., 2025). These measures are not comparisons and ratings of the same stimuli, so they do not settle the question; they show that "compare instead of rate" is a hypothesis about a task, not a law. Before 2026 we found no study that compared pairwise comparisons, galleries, sliders, rankings, and ratings on the same design task with real users; the first head-to-head comparisons of feedback formats appeared in 2026 (Section 32.4). The case for comparisons is strongest where it matters most for this book: when there is no external unit to anchor a rating, and when the context drifts during a session (inference).
Sources cited in Section 16.1 13
- Miller (1956) The magical number seven, plus or minus two: Some limits on our capacity for processing information
- Thurstone (1927) A Law of Comparative Judgment
- Chang et al. (2017) Sequential effects in preference decision: Prior preference assimilates current preference
- Vinson et al. (2019) Decision contamination in the wild: Sequential dependencies in online review ratings
- Kramer and Cartledge (2026) Sequential effects in facial attractiveness judgements: No evidence of stable individual differences
- Hollingworth (1910) The Central Tendency of Judgment
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Koyama and Igarashi (2018) Computational Design with Crowds
- Clarke et al. (2021) Assessing Top- Preferences
- Shah et al. (2014) When is it Better to Compare than to Score?
- Budish and Kessler (2022) Can Market Participants Report Their Preferences Accurately (Enough)?
- Xie et al. (2022) Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation
- Bagaïni et al. (2025) A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures
16.2 Psychophysics #
If judgments are noisy, can the noise itself be used to measure something? That was the program of nineteenth-century psychophysics, the study of how physical stimuli map to sensations, and its answer shaped every model in this chapter.
The starting observation is that a small enough difference cannot be told apart reliably. Ask someone to lift two weights and say which is heavier, and the smallest difference they notice, the just-noticeable difference, grows with the weights themselves: roughly in proportion, so that the ratio of the just-noticeable difference to the weight stays about constant. Fechner named this regularity Weber's law, after Ernst Heinrich Weber's experiments on lifted weights, and built on it (Fechner, 1860). If every just-noticeable difference is one equal step of sensation, then counting steps up from the threshold gives a sensation that grows with the logarithm of the intensity, which is Fechner's law. Fechner also systematized the methods for measuring discrimination, among them the one now called the method of constant stimuli: present a fixed standard with a comparison stimulus many times and record how often the comparison is judged larger.
The method of constant stimuli produces the curve that this chapter is about. Plot the proportion of "comparison is heavier" answers against the true difference, and it rises smoothly from near 0, through one half where the two are equal, to near 1. This S-shaped curve is the psychometric function. Its steepness measures the observer's noise: a precise observer has a steep curve, a noisy one a shallow curve. The just-noticeable difference is usually defined from it, as the difference that is judged correctly on some fixed fraction of trials, often 75%.
A different method gives a different law. Instead of asking people to compare, Stevens (1957) asked them to assign numbers directly, "if this sound is 10, how loud is that one?", a method called magnitude estimation. The numbers grew as a power of the intensity rather than as its logarithm, with an exponent that depends on the sensory continuum. Both laws describe data well within their own method. The lesson for this book is that a number a person gives directly is not a neutral readout of their sensation: it passes through their own mapping from sensation to numbers, while discrimination data measure something else, how often two things are confused. Comparisons measure utilities in units of the person's noise, as Fechner counted sensation in just-noticeable differences. Section 18.4 returns to what this means for a learned utility.
Modern psychophysics has refined the picture without overturning it. Part of the noise and bias in judgments of value is now explained as efficient coding: the brain adapts its scale to the range of values it has recently encountered, so the same option can be valued differently in a different context (Bavard et al., 2018). Section 37.3 and Section 39.4 report this work in detail; here it is one more reason to expect a person's scale to move during a session.
Sources cited in Section 16.2 3
- Fechner (1860) Elemente der Psychophysik
- Stevens (1957) On the Psychophysical Law
- Bavard et al. (2018) Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences
16.3 Thurstone's comparative judgment #
Fechner's psychophysics needed a physical scale, grams or decibels, on which to measure the stimulus. Thurstone (1927) removed that requirement. His law of comparative judgment applies, in his words, "not only to the comparison of physical stimulus intensities but also to qualitative comparative judgments such as those of excellence of specimens", such as handwriting samples, children's drawings, or opinions on public issues. This is the step that makes comparisons useful for preferences: there is no physical scale of how much someone likes a color, but there can still be a psychological one.
Thurstone's model has three ingredients. Each time an observer looks at a stimulus, it evokes a discriminal process, a value on a psychological scale. The process fluctuates from occasion to occasion, which is why the observer gives different answers to the same pair on different occasions. Its most frequent value is the stimulus's scale value , and the standard deviation of its fluctuation is the stimulus's discriminal dispersion . On each occasion the observer reports as better the stimulus whose process is higher at that moment.
Thurstone defined the scale so that the fluctuations are normally distributed, which makes the probability of each answer computable. Write and for the processes evoked by stimuli and on one occasion.
Let and be jointly Gaussian with mean zero, standard deviations and , and correlation .
- The observer answers "" when , that is, when the discriminal difference is positive.
- is a linear function of a Gaussian vector, so it is Gaussian (Section 4.3). Its mean is .
- Its variance is , by the rule for the variance of a difference (Section 2.6).
- Standardize: , and the standardized variable is .
- By the symmetry of the standard normal, , where is the standard normal cumulative distribution function. Hence, writing for " is preferred to ", .
Writing for the observed proportion of "" answers and for the corresponding standard normal deviate, the result is Thurstone's law of comparative judgment:
The correlation is where the shared shifts of Section 16.1 live. A mood that raises the impression of both specimens makes their fluctuations move together, , and the variance of the difference shrinks: the shared part has cancelled. Thurstone also noted the opposite case. In simultaneous contrast, seeing one gray next to a darker one makes it look lighter and the darker one look darker, so the fluctuations move apart, , and the difference is exaggerated (Thurstone, 1927).
Equation Equation (16.1) has too many unknowns to fit in general, so Thurstone listed five cases with progressively stronger assumptions. The last and simplest, Case V, assumes that all discriminal dispersions are equal and the correlation is zero (Thurstone suggested it was "legitimate for rough measurement"). With a common dispersion , the probability becomes
This is the probit choice model, named after the probability unit, an old name for a standard normal deviate. Read it as a recipe: the probability of choosing depends only on the difference in scale values, measured in units of the noise, and passes through one half when the difference is zero. Replace the scale value by a utility function of a design , and Equation (16.2) is the likelihood of Chu and Ghahramani's preference model (Chu and Ghahramani, 2005), used in Chapter 18 and in every chapter after it.
Thurstone used the model in the other direction, to measure. Taking as the unit, Case V gives , his equation (4). If 75% of judgments prefer to , then and sits about above . If 99% prefer , the gap is about . Collecting such gaps for many pairs places every stimulus on one scale, with no physical measurement anywhere.
Suppose is preferred to in 69% of judgments and to in 84%. Case V gives and . If the model holds, , so it predicts that is preferred to in of judgments. Observing the third proportion is therefore a test: the gaps must add up. A large mismatch means the options do not lie on one scale with equal noise, which is the subject of Section 16.7.
Thurstone also observed that proportions are not equally informative. A proportion of 0.99 and one of 0.55 do not pin down their scale differences equally well (Thurstone, 1927): near 0 or 1, a small change in the proportion corresponds to a large change in the difference, so a pair whose outcome is nearly certain says little about how far apart the options are. Section 16.6 makes this precise.
Sources cited in Section 16.3 2
- Thurstone (1927) A Law of Comparative Judgment
- Chu and Ghahramani (2005) Preference learning with Gaussian processes
16.4 Bradley, Terry, and Luce #
A quarter century later, statisticians working on paired-comparison experiments arrived at a second model by a different route. Bradley and Terry (1952) gave each option a positive worth and set
Writing the worth as an exponential of a utility, with a positive scale , and dividing through by turns Equation (16.3) into
the logistic function, , of the scaled utility difference, also called the logit model. Like Case V, it depends only on the difference of utilities, it is one half at a tie, and it approaches 0 and 1 for large differences. The scale plays the role of the noise: a small makes the choice nearly deterministic.
Luce (1959) extended the model from pairs to sets. His choice axiom implies that the probability of choosing from any set of options is
for positive weights , the rule now usually called the softmax when . For a set of two it is the Bradley-Terry model. The same idea gives a model for rankings: choose the first-ranked option from the full set by Equation (16.5), the second from the remaining options, and so on. Plackett (1975) developed this model for permutations, and it is known as the Plackett-Luce model. Section 20.1.3 uses it when a person orders several options instead of answering a duel, the field's word for one pairwise comparison, which the rest of the book uses too.
The axiom has a strong consequence. Divide Equation (16.5) for two options and in the same set: , whatever else is in . This property is the independence of irrelevant alternatives: the odds of choosing tea over coffee do not depend on whether juice is also on offer. It is convenient and often false. A thought experiment shows why: if a café offers coffee and tea and a person picks each half the time, adding a second, identical pot of coffee should not change much, yet under Equation (16.5) the two coffees together take two thirds of the choices.
Real choices violate the property in subtler ways. Adding a third option can make one of the original two more attractive (the attraction or decoy effect), make the middle option more attractive (the compromise effect), or draw share mostly from the option most similar to it (the similarity effect). These context effects hold on average in groups, but individuals rarely show all three, and averaging can produce a pattern that no individual shows (Liew et al., 2016). A dominated option can even make the option that dominates it look worse, a repulsion effect that depends on how the options are displayed (Spektor et al., 2018). A review attributes the appearance, disappearance, and reversal of context effects to the spatial layout of the options, how concrete their attributes are, and how long people deliberate (Spektor et al., 2021). A pure duel has no third option, so the classic effects need a set to act on; galleries, and options still in memory from earlier queries, can bring them back (inference). Section 37.2 reports this evidence, and Section 20.6 what it means for interface design.
Sources cited in Section 16.4 6
- Bradley and Terry (1952) Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
- Luce (1959) Individual Choice Behavior: A Theoretical Analysis
- Plackett (1975) The Analysis of Permutations
- Liew et al. (2016) The appropriacy of averaging in the study of context effects
- Spektor et al. (2018) When the Good Looks Bad: An Experimental Exploration of the Repulsion Effect
- Spektor et al. (2021) The elusiveness of context effects in decision making
16.5 Random utility models #
The probit and the logit look like two unrelated formulas. They are two instances of one idea, which economists developed into the main framework for analyzing choices: the random utility model. Each option has a systematic utility , and on each occasion the person perceives , with random noise ; they choose the option whose perceived utility is largest. Thurstone's discriminal process is a random utility with Gaussian noise. McFadden (1974) showed that when the noise terms are independent with a Gumbel distribution (also called the double exponential or type I extreme value distribution), the choice probabilities are exactly Luce's rule Equation (16.5) with , the conditional logit model, which became the starting point of discrete choice analysis in economics (Section 40.9).
So the choice model is a claim about the noise: Gaussian noise gives the probit, Gumbel noise gives the logit. The curve that turns a utility difference into a choice probability is also called the link function, or link for short, and from here on we speak of the probit link and the logit link. Yellott (1977) studied how Luce's axiom, Thurstone's theory, and the double exponential distribution are connected. For two options the connection is short enough to derive.
Let and be independent Gumbel variables with scale . A location shift common to both cancels in the difference, so take the standard form with cumulative distribution function and density . Let .
- is chosen when . Conditioning on and averaging over (the law of total probability), .
- Substitute . Then , and , since . As runs from to , runs from to .
- The integral becomes .
- The integral of over is , so , the logistic Equation (16.4) with .
The figure below makes the random utility story concrete. The top panel shows the perceived utilities of two options with a given difference and noise. Each press of play simulates one occasion: the person perceives one value for each option, marked by triangles, and chooses the higher. The middle panel shows the probability of choosing against the utility difference, with the share of choices among the occasions so far.
Some things to try:
- Shrink the noise. As falls the curve steepens toward a step: a noise-free person always picks the better option, and the probability carries no information about how much better it is. As grows the curve flattens toward one half everywhere.
- Switch the noise. The figure matches the two noise distributions in standard deviation, which for Gumbel noise means a scale . The two curves then nearly coincide. Their largest gap is about 0.023, at a difference of about 0.68 standard deviations of the perceived difference .
- Look at the tails. The difference shows up far from zero. When the utility difference is 3 standard deviations of , the logit gives the worse option a probability of about 0.0043 and the probit about 0.0013, 3.2 times less. At 4 standard deviations the ratio is about 22.
- Turn on the cost. The penalty a model pays, in log likelihood, for an answer against the difference grows linearly in the difference for the logit and quadratically for the probit. One surprising answer pulls a probit model much harder than a logit model.
- Press play. Each occasion is a fresh draw. With 60 occasions the share of choices usually lands within its interval of the curve, but any single answer can go either way.
The last two points matter in practice. A person who is careless once, or misreads one pair, gives an answer far out in the tail. Under the probit that answer can dominate the fit; under the logit its influence is bounded. BoTorch's probit implementation clips the argument of to , which caps the penalty for any single comparison, a safeguard that Section 27.5 examines (Meta Platforms, Inc., 2026g).
Papers write the probit link in at least three ways:
with noise on each option, as in Equation (16.2);
with noise on the difference; and with the noise
absorbed into the scale of . BoTorch's PairwiseGP uses
, which is per-option noise fixed at 1
(Meta Platforms, Inc., 2026g). The three describe the same model with different
units, but a noise value copied from one paper into another's formula is off by
a factor of . The same care applies when comparing the probit with
the logit: match standard deviations, , not the raw
parameters.
What is the noise? A random utility model is agnostic. The randomness can be a person's moment-to-moment fluctuation, as in Thurstone's single observer; it can be attributes of the options that the analyst does not observe; or it can be variation across the people in a sample. For one person answering a sequence of comparisons, the first reading is the natural one, and it can be tested. McCausland et al. (2020) asked 141 participants to choose among five lotteries, six times from every subset of at least two, and applied a set of inequalities that choice probabilities must satisfy if any random utility model generated them. Most participants were consistent with random utility; only 4 showed strong evidence of violating it.
One more property is shared by every model in this section, and it shapes the rest of Part IV. The choice probability depends on utilities only through (or ). Adding a constant to every utility changes nothing, and doubling every utility while doubling the noise changes nothing. Comparisons can therefore determine utilities only up to a shift, and only in units of the noise. Section 18.4 works out what this means for a Gaussian process utility.
Probit and logit are the same model, a random utility with the better-looking option chosen, under Gaussian and Gumbel noise. They agree near a tie and differ in the tails, where they disagree about how surprising a surprising answer is.
Sources cited in Section 16.5 4
- McFadden (1974) Conditional Logit Analysis of Qualitative Choice Behavior
- Yellott (1977) The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution
- Meta Platforms, Inc. (2026g) BoTorch pairwise likelihood source code likelihoods/pairwise.py
- McCausland et al. (2020) Testing the Random Utility Hypothesis Directly
16.6 What a comparison carries #
A rating on a scale of 1 to 9 can, in principle, carry bits. A comparison has two possible answers, so it can carry at most one bit: the information an answer gives about anything is bounded by the entropy of the answer, and the entropy of a binary answer is at most bit (Section 6.3). Most comparisons carry much less. To see how much, and which comparisons carry the most, combine the choice model with what the model already believes.
Suppose the model's belief about the difference is Gaussian, : its best guess is and its uncertainty is . The answer follows the probit Equation (16.2) with per-option noise ; write for the noise on the difference. Before asking, the model predicts the answer by averaging the choice probability over its belief.
- Write the choice probability as an event: for an independent , by the definition of .
- Average over the belief: , by the law of total probability.
- is a sum of independent Gaussians, so it is Gaussian with mean and variance (Section 4.6).
- Standardizing as in the derivation of Equation (16.2), .
The model's own uncertainty adds to the person's noise. A pair the model is unsure about is predicted closer to one half than the same pair would be if the model knew the difference. Section 18.3 uses exactly this formula to predict a new comparison from a Gaussian process posterior.
How much will the answer teach? The answer is uncertain for two reasons: the model does not know , and even if it did, the person is noisy. Only the first kind of uncertainty can be reduced by asking, so the information the answer carries about is the total uncertainty minus the noise part:
where is the entropy, in bits, of a yes-or-no answer with probability , and the expectation is over the belief . The first term is the entropy of the predicted answer; the second is the entropy the answer would keep if were known, averaged over the values the model considers plausible. This decomposition is the basis of the Bayesian active learning by disagreement criterion of Houlsby et al. (2011), which they also applied to preference learning: the most informative question is one whose answer the model cannot predict, but whose answer would be predictable if the model knew the truth.
The figure shows why two kinds of pairs teach little. Some things to try:
- A known near-tie. Set the mean to zero and shrink the belief's spread . The predicted answer is a coin flip, one full bit of uncertainty, but almost all of it is the person's noise. The answer carries almost nothing, 0.006 bits at and , because the model already knows the options are nearly equal.
- A known gap. Move the mean far from zero. The answer is foregone, both entropies fall toward zero, and so does the information: about 0.02 bits at with and .
- An open question. With the mean near zero and the spread large compared with the noise, the model cannot predict the answer and would be able to if it knew . This is where a comparison is worth asking: about 0.6 bits at , , , rising toward a full bit as the noise vanishes.
- A noisy person. Raise . The dotted curve rises toward the dashed one and every comparison teaches less. Noise cannot be designed away; it sets the price of each answer.
The most useful comparison, then, pairs options whose order the model is unsure of, relative to how noisy the person is. That is the intuition behind the acquisition rules of Section 19.3 and Section 19.4, and behind the query designs of Chapter 20.
Fewer bits per answer do not necessarily mean slower learning. For estimating the utilities of a fixed set of options under the Thurstone and Bradley-Terry models, Shah et al. (2016) proved minimax bounds, the best error any method can guarantee in the worst case, and found that the error depends on the topology of the comparison graph (which pairs were compared) through the eigenvalues of its graph Laplacian, a matrix that records which pairs were compared, and that the ordinal and cardinal settings have error rates with the same scaling, up to constant factors. Each comparison carries less than a numerical measurement would, but the rate at which errors shrink with more data is the same (inference). Section 18.5 explains what the comparison graph is and why its structure matters for a Gaussian process utility too.
A related result tempers the role of the link function. For actively ranking a set of items from noisy comparisons, Heckel et al. (2019) showed that a simple counting method that assumes no parametric model is optimal up to logarithmic factors, so parametric assumptions such as Bradley-Terry or Thurstone buy at most a logarithmic gain. For preferential Bayesian optimization (PBO) this suggests that sample efficiency comes mainly from the kernel sharing information between nearby inputs, not from the exact form of the link (inference).
The point sharpens as the number of design parameters grows. A comparison between two exoskeleton gaits that differ in four parameters (Li et al., 2021), or between two simplifications of a 3D mesh controlled by nine parameters (Ou et al., 2022), still returns one bit at most, while the number of designs that would have to be told apart grows exponentially with the number of parameters. Sessions with people rarely run beyond a few dozen comparisons (Part VII), so no choice model can extract more than a few dozen bits from one. The rest has to come from assumptions about how utility varies across the design space, the kernel of Chapter 18, and from the choice of which pairs to ask, the subject of Chapter 19 (inference).
Sources cited in Section 16.6 5
- Houlsby et al. (2011) Bayesian Active Learning for Classification and Preference Learning
- Shah et al. (2016) Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
- Heckel et al. (2019) Active ranking from pairwise comparisons and when parametric assumptions do not help
- Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
16.7 Assumptions to watch #
Every model in this chapter, and the Gaussian process preference model built on them in Chapter 18, makes assumptions about the person answering. They are reasonable starting points, and each has been tested. Table 16.1 lists them with the evidence; the paragraphs after it add what the table cannot hold.
| Assumption | What the evidence says | Where the book returns |
|---|---|---|
| One stable utility behind every answer | Choosing changes preferences: a meta-analysis of 43 artifact-free studies finds a shift of standard deviations | Section 37.2, Section 46.5 |
| Answers are independent given the utility | Choices and ratings depend on the preceding trials | Section 37.3 |
| Preferences are transitive | Most people satisfy random utility; true cycles exist in specific designs | Section 37.4, Chapter 21 |
| The noise is the same for every pair | How noise is specified changes the inferred preferences | Section 37.4, Section 27.2 |
| Every forced choice reflects a preference | People report preferences even between identical samples | Section 20.4 |
| Choices do not depend on other options | Context effects exist but are conditional | Section 37.2, Section 20.1 |
| The utility scale is fixed across sessions | Values adapt to the recent range | Section 37.3 |
Stability. The models treat the person as a fixed utility plus noise. But choosing can change what people like. A meta-analysis of 43 studies using the free-choice paradigm with the known artifact removed (N = 2,191) found that after people choose between two similar options, they rate the chosen one higher and the rejected one lower, with an effect size of , a shift of 0.40 standard deviations (95% confidence interval 0.32 to 0.49), and no evidence of publication bias (Enisman et al., 2021). A sequential-sampling account explains part of the mechanism: each choice raises the value of the chosen option and lowers that of the rejected one, and the consistency of repeated choices between the same pair declines as more trials intervene (Zylberberg et al., 2024). An optimizer that keeps showing its current favorite may therefore be reinforcing it (inference).
Independence. The likelihood multiplies the probabilities of the answers as if each were a fresh draw. The serial dependence just described, and the sequential effects on ratings in Section 16.1, say that an answer partly depends on what came before. The effects are reliable on average but unstable within individuals (Kramer and Cartledge, 2026), which makes them hard to correct one person at a time.
Transitivity. Random utility models with independent noise are transitive in a probabilistic sense: if usually beats and usually beats , then usually beats . Direct tests mostly support this (McCausland et al., 2020), but not all observed cycles are noise. Using response times to separate noise from preference, a 2023 working paper found that transitivity violations shrink but do not disappear: on average across participants, 19.24% and 13.83% of the cycles with revealed preferences in two reanalyzed data sets were violations, most often arising from chains of small trade-offs between attributes (Alós-Ferrer et al., 2023). Lotteries designed after the Steinhaus-Trybula paradox likewise produced cycles as the most common pattern even after allowing for transitive preferences with noise (Butler and Pogrebna, 2018). Chapter 21 discusses what can be optimized when no utility exists.
Homogeneous noise. Case V and the logit give every pair the same noise. In risky choice, how the noise is specified changes what is inferred: combining noise in preferences with noise in responding can make an expected-value maximizer look risk averse or risk seeking (Bhatia and Loomes, 2017), and with homogeneous noise the inferred risk aversion can behave non-monotonically, which led Apesteguia and Ballester (2018) to recommend random-parameter models instead. Section 27.2 lists the preference models that let noise vary.
Forced choice. The duel offers no "neither" and no "I can't tell". In sensory science, consumers asked to choose between two identical samples still report a preference, which is why that field moved to "no preference" options and placebo pairs (O'Mahony and Wichchukit, 2017). Section 20.4 covers likelihoods that allow ties.
Context and scale. The models assume the choice between two options does not depend on what else has been seen, and that the utility scale is the same in every session. The context effects of Section 16.4 and the range adaptation of Section 16.2 both cut against this. If a person's values are normalized to the range seen in a session, a utility learned in one session is on a session-specific scale, and reusing it in another session needs recalibration (inference).
None of this means the models are wrong to use. It means their output is an estimate under assumptions that can be checked: by repeating a few early pairs late in a session, by allowing ties, by randomizing which side an option appears on, and by comparing the fitted noise level with the observed rate of reversed answers. Part IX asks the deeper question these checks circle around, whether repeated comparisons find a preference or partly make one (Chapter 45).
Sources cited in Section 16.7 9
- Enisman et al. (2021) Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm
- Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
- Kramer and Cartledge (2026) Sequential effects in facial attractiveness judgements: No evidence of stable individual differences
- McCausland et al. (2020) Testing the Random Utility Hypothesis Directly
- Alós-Ferrer et al. (2023) Identifying Nontransitive Preferences
- Butler and Pogrebna (2018) Predictably intransitive preferences
- Bhatia and Loomes (2017) Noisy preferences in risky choice: A cautionary note
- Apesteguia and Ballester (2018) Monotone Stochastic Choice Models: The Case of Risk and Time Preferences
- O'Mahony and Wichchukit (2017) The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection
16.8 Exercises #
In a taste test, beats in 75% of trials, beats in 75%, and beats in 84%. Under Case V, place the three options on one scale with , and check whether the three proportions are consistent with it. What would a logistic model with matched standard deviation predict for against ?
Solution
From Equation (16.2), . If the scale is consistent, , and Case V predicts . The observed 84% is lower, so the gaps do not add up: either the noise differs between pairs (Cases I to IV allow this), or the proportions are noisy estimates. With matched standard deviation the logistic scale is , so against has probability , almost the same as the probit. Near the middle of the curve the two links are nearly indistinguishable, and a test like this one cannot tell them apart.
Extend the derivation of the logistic link to options. Let with independent standard Gumbel noise of scale . Show that the probability that option has the largest perceived utility is , Luce's rule Equation (16.5).
Solution
Condition on . Option 1 wins when for every , which by independence has probability . Substitute as before, so that , and write . The probability is . Multiplying numerator and denominator by gives .
Using Equation (16.7), show that the information carried by one comparison tends to zero (a) as for any fixed and , and (b) as for fixed and . Then show that it tends to one bit when and with fixed.
Solution
(a) As the belief concentrates at , so the expectation in the second term tends to , and the first term tends to the same value, since . The difference tends to zero. (b) As grows, tends to 0 or 1, so the first term tends to zero; the second term is nonnegative and no larger than the first, because information is never negative, so it tends to zero too. (c) With the first term is bit for every . As , tends to 0 or 1 for every , so the second term tends to zero, and the information tends to one bit: a noise-free answer about a difference whose sign is a fair coin is worth exactly one bit.
Further reading #
- Thurstone (1927) is short and readable; it defines the discriminal process, states the law, and lists the five cases.
- Bradley and Terry (1952) introduce the paired-comparison model now named after them; Luce (1959) develops the choice axiom and its consequences; Plackett (1975) extends it to rankings.
- McFadden (1974) derives the conditional logit from random utility with Gumbel noise; Yellott (1977) connects Luce, Thurstone, and the double exponential distribution.
- Miller (1956) is the classic account of how little information an absolute judgment carries, and why relative judgments help.
- Fechner (1860) and Stevens (1957) are the two sides of the debate over measuring sensation by discrimination or by direct numbers.
- Shah et al. (2016) and Shah et al. (2014) compare ordinal and cardinal measurement, in theory and in crowdsourcing experiments.
- Houlsby et al. (2011) derive the information criterion of Section 16.6 and apply it to preference learning.
References
- (2023). Identifying Nontransitive Preferences. University of Zurich. working paper Cited in §16.7
- (2018). Monotone Stochastic Choice Models: The Case of Risk and Time Preferences. Journal of Political Economy. Cited in §16.7
- (2025). A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures. Nature Human Behaviour. doi:10.1038/s41562-024-02085-2. Cited in §16.1
- (2018). Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences. Nature Communications. Cited in §16.2
- (2017). Noisy preferences in risky choice: A cautionary note. Psychological Review. Cited in §16.7
- (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika. Cited in §16.4
- (2022). Can Market Participants Report Their Preferences Accurately (Enough)? Management Science. Cited in §16.1
- (2018). Predictably intransitive preferences. Judgment and Decision Making. Cited in §16.7
- (2017). Sequential effects in preference decision: Prior preference assimilates current preference. PLOS ONE. Cited in §16.1
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Cited in §16.3
- (2021). Assessing Top- Preferences. ACM Transactions on Information Systems. Cited in §16.1
- (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Cited in §16.7
- (1860). Elemente der Psychophysik. Breitkopf und Härtel. Cited in §16.2
- (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. Cited in §16.6
- (1910). The Central Tendency of Judgment. The Journal of Philosophy, Psychology and Scientific Methods. Cited in §16.1
- (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Cited in §16.6
- (2018). Computational Design with Crowds. Computational Interaction. Cited in §16.1
- (2026). Sequential effects in facial attractiveness judgements: No evidence of stable individual differences. Perception. Cited in §16.1 §16.7
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Cited in §16.6
- (2016). The appropriacy of averaging in the study of context effects. Psychonomic Bulletin & Review. Cited in §16.4
- (1959). Individual Choice Behavior: A Theoretical Analysis. Wiley. Cited in §16.4
- (2020). Testing the Random Utility Hypothesis Directly. The Economic Journal. doi:10.1093/ej/uez039. Cited in §16.5 §16.7
- (1974). Conditional Logit Analysis of Qualitative Choice Behavior. Frontiers in Econometrics. Cited in §16.5
- (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. software Cited in §16.5
- (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review. Cited in §16.1
- (2017). The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection. Trends in Food Science & Technology. Cited in §16.7
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §16.1 §16.6
- (1975). The Analysis of Permutations. Journal of the Royal Statistical Society: Series C (Applied Statistics). Cited in §16.4
- (2014). When is it Better to Compare than to Score? arXiv. preprint Cited in §16.1
- (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Cited in §16.6
- (2018). When the Good Looks Bad: An Experimental Exploration of the Repulsion Effect. Psychological Science. Cited in §16.4
- (2021). The elusiveness of context effects in decision making. Trends in Cognitive Sciences. Cited in §16.4
- (1957). On the Psychophysical Law. Psychological Review. Cited in §16.2
- (1927). A Law of Comparative Judgment. Psychological Review. Cited in §16.1 §16.3
- (2019). Decision contamination in the wild: Sequential dependencies in online review ratings. Behavior Research Methods. Cited in §16.1
- (2022). Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation. Quality of Life Research. Cited in §16.1
- (1977). The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution. Journal of Mathematical Psychology. Cited in §16.5
- (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Cited in §16.7