Bayesian Optimization
Part IX: What Is a Preference?
中文

Natural and Formal Sciences

The previous chapter asked where preferences come from and what happens when a system measures them. This chapter turns to disciplines that model choice with mathematics. What they offer preferential Bayesian optimization (PBO) is an account of what a set of comparisons can and cannot tell a model. The most useful results are tools a PBO system can apply to its own logs today: the shape of the comparison graph decides how well the utilities are pinned down (Section 43.3.1), and a decomposition of the comparison data measures how much of it any single utility can explain (Section 43.3.2). Around them, the chapter checks formal claims that circulate in writing about preference learning, several of them wrong in their sources or conditions (Section 43.9), and ends with what privacy would cost a preference session.

43.1 Biology and behavioral ecology #

Behavioral ecology studies animal choice as an adaptation, which makes it a natural test of whether preferences are stable, scalar, and transitive. Its answer is mostly a transitive core with context and state effects around it. Of 15 honeybees given binary choices between artificial flowers, 3 violated weak stochastic transitivity (if aa beats bb and bb beats cc at least half the time, aa beats cc at least half the time), under an assumption about how the flowers ranked on a utility scale (Shafir, 1994). Slime molds, often offered as an example of intransitive choice, had a linear, transitive ranking of food options in the original study (Latty and Beekman, 2011); they violated only the independence of irrelevant alternatives, the principle that adding a third option should not change the relative preference between two others. Decoy effects (Section 37.2) depend on state and design. In starlings they appear or disappear with the animal's state (Schuck-Paim et al., 2004). In bumblebees, decoys differing in reward rate shifted preferences as predicted, while decoys differing in sugar concentration did not (Hemingway et al., 2024), and a 2025 preprint, later published in Ecological Entomology in 2026, found that adding rewardless flowers, a different decoy design, did not raise preference for neighboring flowers, and very little support for decoy effects in bees in earlier research (Armand et al., 2026).

Two further results bear on a session's dynamics. Harhen and Bornstein (2023) showed that "overharvesting", staying in a food patch longer than Charnov's marginal value theorem prescribes (Charnov, 1976), can follow from rational inference about the environment combined with discounting adjusted for uncertainty, and that human participants behaved in line with their model. And evolutionary theory predicts that the utility scale itself adapts, rising most steeply where choices are frequent and mistakes are costly (Netzer, 2009).

What it means for PBO. A user who "stays too long" near the incumbent may be acting rationally under uncertainty, so a stopping rule should compare the marginal expected gain with the posterior distribution of the attainable level, not with a point estimate (inference; Section 46.6). If utility scales adapt within a session, PBO's concentration of queries near the optimum would sharpen discrimination there, against the assumption of a fixed floor set by Weber's law, under which the smallest noticeable difference grows with the magnitude (Section 37.3) (inference); Section 43.5 describes how to test it. And because the core is transitive, a model should add covariates for the choice set, recent history, and state to the likelihood before switching to a surrogate that can represent intransitive preferences (inference).

Sources cited in Section 43.1 8
  1. Shafir (1994) Intransitivity of preferences in honey bees: support for 'comparative' evaluation of foraging options
  2. Latty and Beekman (2011) Irrational decision-making in an amoeboid organism: transitivity and context-dependent preferences
  3. Schuck-Paim et al. (2004) State-Dependent Decisions Cause Apparent Violations of Rationality in Animal Choice
  4. Hemingway et al. (2024) Economic foraging in a floral marketplace: asymmetrically dominated decoy effects in bumblebees
  5. Armand et al. (2026) No evidence of a decoy effect in bees: Rewardless flowers do not increase bumblebees' preference for neighbouring flowers
  6. Harhen and Bornstein (2023) Overharvesting in human patch foraging reflects rational structure learning and adaptive planning
  7. Charnov (1976) Optimal foraging, the marginal value theorem
  8. Netzer (2009) Evolution of Time Preferences and Attitudes toward Risk

43.2 Physics and the statistical mechanics of decisions #

The claim and the check. In Luce's choice axiom, the probability of choosing option xx from a set SS is w(x)/∑y∈Sw(y)w(x) / \sum_{y \in S} w(y) for positive weights ww (Luce, 1959). Writing w(x)=eβf(x)w(x) = e^{\beta f(x)} makes this the Boltzmann distribution of statistical physics, with the negative utility −f-f as the energy and the inverse temperature β\beta as the precision of choice: at high temperature choices are nearly random, at low temperature nearly deterministic. The identity is correct. The claim built on it, that the PBO likelihood is a softmax and so a Boltzmann distribution, holds only with a logistic link, such as that of González et al. (2017). It does not hold for the probit likelihood of Chu and Ghahramani or for BoTorch's default PairwiseProbitLikelihood (Meta Platforms, Inc., 2026g), which use the Gaussian cumulative distribution function (Section 27.1). Likewise, Luce's axiom and Thurstone's Case V are not equivalent under a logistic distribution, contrary to a claim often made citing Yellott (1977): Luce's form follows when the errors of a random utility model are Gumbel (double-exponential), so that their differences are logistic, while Case V assumes normal errors (Section 16.5).

The precision is not fixed. Lindig-León et al. (2022) found in a two-alternative forced-choice task that participants relied on several stimulus features when they had time and decided mainly on a single feature under high time pressure. Rational inattention, the economic theory in which attention is a costly resource (Section 40.5), arrives at a generalization: Matějka and McKay (2015) proved that optimal acquisition of information yields a generalized multinomial logit model in which choice probabilities depend on both the options' true payoffs and the decision maker's prior beliefs.

What it means for PBO. Rational inattention gives choice probabilities proportional to a prior weight times eβ⋅utilitye^{\beta \cdot \text{utility}}. For a pair of options this reads

P(x chosen over x′)=w(x) eβf(x)w(x) eβf(x)+w(x′) eβf(x′)=sigmoid⁡ ⁣(β(f(x)−f(x′))+log⁡w(x)w(x′)),\Prob(\vx \text{ chosen over } \vx') = \frac{w(\vx)\, e^{\beta f(\vx)}}{w(\vx)\, e^{\beta f(\vx)} + w(\vx')\, e^{\beta f(\vx')}} = \operatorname{sigmoid}\!\Big(\beta\big(f(\vx) - f(\vx')\big) + \log\frac{w(\vx)}{w(\vx')}\Big),
(43.1)

and the standard logistic link is the special case of equal weights. A PBO likelihood could therefore add a default or familiarity term log⁡w\log w for the incumbent or for the option shown first (inference; the abstract of Matějka and McKay confirms only that the choice probabilities depend on the prior, and the specific form above is a summary). Because under time pressure the effective utility may collapse onto a single feature, changing an interface's response deadline changes the structure of the revealed utility, not only the noise (inference). And the temperature describes the user, so it should be estimated, not annealed over the iterations as is sometimes proposed; annealing belongs to the acquisition function, for example the temperature of Thompson sampling (Section 12.5) (inference).

Sources cited in Section 43.2 6
  1. Luce (1959) Individual Choice Behavior: A Theoretical Analysis
  2. González et al. (2017) Preferential Bayesian Optimization
  3. Meta Platforms, Inc. (2026g) BoTorch pairwise likelihood source code likelihoods/pairwise.py
  4. Yellott (1977) The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution
  5. Lindig-León et al. (2022) From Bayes-optimal to heuristic decision-making in a two-alternative forced choice task with an information-theoretic bounded rationality model
  6. Matějka and McKay (2015) Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model

43.3 Network science: the comparison graph #

Network science's most useful contribution to PBO since 2017 is the theory of the comparison graph, whose nodes are the compared options and whose edges are the answered comparisons. Section 27.6 showed that the likelihood Hessian of the Laplace approximation is this graph's Laplacian matrix; this section asks how the graph's shape governs how well the utilities are estimated, and how much of the data one utility can explain.

43.3.1 Comparison graphs and estimation error #

Spectral graph theory studies a graph through the eigenvalues of matrices built from it, most often the Laplacian.

Definition 43.1 Graph Laplacian, algebraic connectivity, effective resistance

For a comparison graph on nn options, the graph Laplacian L\mL is the sum of (ei−ej)(ei−ej)⊤(\mathbf{e}_i - \mathbf{e}_j)(\mathbf{e}_i - \mathbf{e}_j)^\T over the answered pairs (i,j)(i, j), where ei\mathbf{e}_i is the ii-th unit vector. Its eigenvalues are 0=λ1≤λ2≤⋯≤λn0 = \lambda_1 \le \lambda_2 \le \dots \le \lambda_n, and the number of zero eigenvalues equals the number of connected components.

  • The algebraic connectivity is λ2\lambda_2. It is positive exactly when the graph is connected, and the larger it is, the harder the graph is to cut into two weakly linked halves (Fiedler, 1973).
  • The effective resistance RijR_{ij} between options ii and jj is the electrical resistance between them when every comparison is a resistor of one ohm: Rij=(ei−ej)⊤L+(ei−ej)R_{ij} = (\mathbf{e}_i - \mathbf{e}_j)^\T \mL^{+} (\mathbf{e}_i - \mathbf{e}_j), where L+\mL^{+} is the pseudo-inverse of L\mL.

The effective resistance has a direct statistical reading.

Derivation Why effective resistance is the variance of a utility difference
  1. Suppose each answered pair (i,j)(i, j) yields a noisy measurement of the utility difference si−sjs_i - s_j with unit precision: the Gaussian picture behind the Laplace approximation of Section 18.2, with unit curvature per comparison and no prior. The precision (inverse covariance) of the least-squares estimate is then the sum of all the pairs' contributions, which is L\mL by Definition 43.1.
  2. L\mL is singular, because adding a constant to every utility changes no difference, but differences within a connected component have covariance given by L+\mL^{+}. So the variance of the estimated difference between ii and jj is RijR_{ij}, and it is infinite if no path of comparisons links them.
  3. The rules for resistors become rules for query design. Comparisons in series add: the ends of a chain of kk comparisons have R=kR = k. Comparisons in parallel combine like parallel resistors: asking a pair twice halves its resistance.

Three results make this picture quantitative. Shah et al. proved minimax bounds for estimating the utilities under the Bradley-Terry and Thurstone models, showing that the error depends on the topology of the comparison graph through its Laplacian spectrum, and that the error rates for ordinal data (comparisons) and cardinal data (numerical ratings) are identical up to constant factors (Shah et al., 2016). Hendrickx et al. (2019) proved that, when every compared pair is asked many times, the relative error of the estimates scales with the square root of the effective resistance, with a matching lower bound up to logarithmic factors. And Heckel et al. (2019) proved that, for actively ranking options from noisy comparisons, a simple counting algorithm that assumes no model is optimal up to logarithmic factors: parametric assumptions such as Bradley-Terry or Thurstone buy at most a logarithmic improvement.

Figure 43.1 makes the derivation concrete. It arranges twelve compared designs on a circle and lets you choose the order in which a session adds comparisons.

compared designs and answered pairsincumbent0.71∞∞∞∞∞∞∞∞∞∞sd of utility difference to the incumbent01.6not tiedalgebraic connectivity λ₂ of the comparison graph0.00.51.01.52.0λ₂05101520comparisons answeredDisjoint pairsIncumbent vs challengerChainRandom pairs11 comparisons · 6 components · λ₂ = 0.0010 designs not tied to the incumbent at all
compared designs and answered pairsincumbentsd of utility difference to the incumbent01.6not tiedalgebraic connectivity λ₂ of the comparison graph0.00.51.01.52.0λ₂05101520comparisons answeredDisjoint pairsIncumbent vs challengerChainRandom pairs11 comparisons · 6 components · λ₂ = 0.0010 designs not tied to the incumbent at all
Figure 43.1 How comparisons tie utilities together. Left: twelve compared designs, with the incumbent at the top; lines are answered comparisons, and the most recent one is orange. Each design is shaded by the standard deviation of its utility difference to the incumbent: with no prior this is the square root of the effective resistance, and a dashed circle means no chain of comparisons links the design to the incumbent at all. Right: the algebraic connectivity λ₂ of the comparison graph as comparisons are added, for all four query designs. Disjoint pairs mimics queries that never share an input, as EUBO tends to choose. With a Gaussian process prior, the designs are random points in d dimensions with an RBF kernel of lengthscale 0.5, close to the mode of BoTorch's default prior. Each comparison has unit curvature; the values are illustrative.

Things to try:

  • With Disjoint pairs at 11 comparisons, the graph has 6 components and λ2=0\lambda_2 = 0: ten designs are not tied to the incumbent at all. Drag to 24 comparisons. The same pairs are asked again, each pair becomes tighter, and λ2\lambda_2 stays at zero.
  • Switch to Incumbent vs challenger. The graph connects at the 11th comparison, λ2\lambda_2 jumps to 1, and every design's difference to the incumbent has standard deviation 1.00; at 22 comparisons λ2=2\lambda_2 = 2 and the standard deviation falls to 0.71.
  • Switch to Chain. It also connects at 11 comparisons, but λ2\lambda_2 is only 0.07 and the design at the far end has standard deviation 3.32, the square root of 11 resistors in series. Connected is not the same as well tied.
  • Set the prior to GP, d = 1 with Disjoint pairs. The kernel ties nearby designs together and the worst standard deviation is 0.68. Now raise the dimension: with d=10d = 10 it is 1.09, and with d=30d = 30 it is 1.13. In high dimension, random designs sit far apart relative to the lengthscale, the kernel ties almost nothing, and the shape of the comparison graph again decides what is known (compare Section 30.1).

What it means for PBO. The following are inferences. A PBO acquisition function is, in effect, designing a comparison graph on the evaluated points, and the kernel's correlations act as extra edges. The algebraic connectivity, or the effective resistance, among the evaluated points is a practical diagnostic: regions joined to the rest by few comparisons have poorly anchored utilities, and if the kernel is misspecified, the posterior variance will not show it. Heckel et al.'s result implies that a surrogate gains mainly from the kernel's smoothness and little from the choice of link (Section 27.1).

43.3.2 Cycles and the Hodge decomposition #

A scalar utility can only produce comparisons that are consistent around every loop: if AA beats BB by a margin and BB beats CC by another, the utility fixes the margin of AA over CC as their sum. Real comparison data are rarely that consistent, and the Hodge decomposition of comparison data (Jiang et al., 2011), a foundational result from before the period this part covers, measures by how much. It treats the comparison margins (for example the log-odds that ii beats jj) as a flow on the edges of the comparison graph and splits the flow into three orthogonal parts:

  • a gradient part, the differences si−sjs_i - s_j of one score per option, found by least squares (this estimate is called HodgeRank); it is the part that one scalar utility explains;
  • a curl part, made of local cycles around triangles of options;
  • a harmonic part, made of global cycles around larger loops in the graph that triangles do not fill in.

Because the parts are orthogonal, the squared size of each part divided by the squared size of the whole flow is the share of the data it accounts for. Since 2017, Strang et al. (2022) used the decomposition to quantify cyclic competition in tournaments, and the skew-symmetric "generalized preference kernel" of Chau et al. (2022) lets a Gaussian process represent cycles (Section 27.2).

Figure 43.2 applies the decomposition to five options with a rock-paper-scissors component of adjustable strength added among three of them, and shows the catch in using it on real data.

answers on every pairABCDEarrow from the option chosen more often;width: |log-odds|; magenta: the cyclic residualmajorities are transitiveshare of the comparison flowexplained by one utility 92%cyclic 8%utility per option−101ABCDEtransitive part of the truthHodgeRank score
answers on every pairABCDEarrow from the option chosen more often;width: |log-odds|; magenta: the cyclic residualmajorities are transitiveshare of the comparison flowexplained by one utility 92%cyclic 8%utility per option−101ABCDEtransitive part of the truthHodgeRank score
Figure 43.2 How much of a set of comparisons one utility can explain. Five options have utilities 1, 0.5, 0, −0.5, −1, plus a rock-paper-scissors flow of strength κ among A, B, and C. Every pair is answered n times under a Bradley-Terry model, and each edge carries the empirical log-odds. The bar splits the flow into its gradient part (blue, explained by one utility) and its cyclic part (magenta); the dashed mark shows the cyclic share that a purely transitive person would produce by chance with the same n, averaged over 200 simulated sessions. Below, the HodgeRank scores against the transitive part of the truth. The setup is illustrative.

Things to try:

  • The figure opens with exact win probabilities. At κ = 0.6 the cyclic share is 8%, the magenta residual sits only on the triangle A, B, C, and the majority preferences are still transitive; the cycle in the majorities (A beats B, B beats C, C beats A) appears only above κ = 1. The HodgeRank scores recover the transitive part exactly, because a pure cycle is orthogonal to every gradient.
  • Choose 100 answers per pair. At κ = 0.6 the first draw shows a cyclic share of 11% against a chance level of 2%: the cycle is detectable.
  • Choose 5 answers per pair. The chance level rises to 27%, and the first draw shows 34%, too close to tell apart; press Draw new answers a few times to see how much the share varies. With 1 answer per pair, the typical PBO situation, a perfectly transitive person already shows a cyclic share of 47% on average.

What it means for PBO. The Hodge decomposition is a cheap intransitivity diagnostic for PBO logs: the share of the curl and harmonic parts in the comparison flow measures the fraction of the data that a scalar Gaussian process utility cannot explain, and only a large share justifies a skew-symmetric surrogate (inference). As the figure shows, the share must be compared with what sampling noise alone produces, so a session that wants to test for intransitivity needs repeated pairs or nearby designs pooled into nodes (inference).

Sources cited in Section 43.3 7
  1. Fiedler (1973) Algebraic Connectivity of Graphs
  2. Shah et al. (2016) Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
  3. Hendrickx et al. (2019) Graph Resistance and Learning from Pairwise Comparisons
  4. Heckel et al. (2019) Active ranking from pairwise comparisons and when parametric assumptions do not help
  5. Jiang et al. (2011) Statistical ranking and combinatorial Hodge theory
  6. Strang et al. (2022) The Network HHD: Quantifying Cyclic Competition in Trait-Performance Models of Tournaments
  7. Chau et al. (2022) Learning Inconsistent Preferences with Gaussian Processes

43.4 The mathematics of preference #

Two questions from the mathematics of order matter for PBO: when a preference can be summarized by a utility function, and what to do when it is not complete.

43.4.1 Representation theorems, checked #

A utility representation of a preference is a function uu with xx weakly preferred to yy exactly when u(x)≥u(y)u(x) \ge u(y). Popular statements of when one exists are often too strong. Debreu's theorem (Debreu, 1964) is commonly stated for any topological space: a continuous representation exists if and only if the preference is complete (any two options can be compared), transitive, and continuous (the sets of options better and worse than any given option are closed). In fact, as the survey of Hervés‐Beloso and del Valle‐Inclán Cruces (2019) explains, separability (a countable dense subset) is necessary, a connected and separable space, or a second-countable one, suffices, and every non-separable metric space carries a continuous preference order with no utility representation at all. The lexicographic order on R2\R^2, which ranks by the first coordinate and uses the second only to break ties, is said to lack a continuous representation; it cannot be represented by any real-valued function, continuous or not (Banerjee and Mitra, 2018).

What it means for PBO. PBO evaluates only finitely many points, and every complete and transitive relation on a countable set has a utility representation, so these failures concern extending a utility to the continuum. A stationary, smooth Gaussian process cannot represent the lexicographic order, but on any finite set of points it can approximate it with a very short lengthscale in the secondary dimension, with input warping, or with threshold features: the practical issue is kernel design, not whether a utility exists (inference; Chapter 9).

43.4.2 Incomplete preferences and contextuality #

Some pairs may simply not be comparable for a person. A multi-utility representation handles this with a set of utility functions, declaring xx better than yy only when every function in the set agrees (Evren and Ok, 2011, a foundational paper). In choice data, a person who picks each of two options about half the time may be indifferent (the options are equally good), indecisive (they cannot rank them), or willing to experiment. Ok and Tserenjigmid (2022) gave methods to identify which, and showed that each identification yields a way to make deterministic welfare comparisons from random choice. The distinction matters in practice. In the experiments of Cettolin and Riedl (2019), about half of the participants chose in a way inconsistent with complete preferences plus certainty independence (an axiom that says mixing two options with the same sure outcome should not change which is preferred). Of those participants, about half behaved consistently with incomplete preferences and about a third with a preference for randomization, and further experiments showed that probability weighting, errors, regret aversion, or intransitive indifference could not explain the pattern.

A stronger claim is that preferences are contextual in the sense of quantum physics: no single set of underlying values explains the measurements, even after allowing each measurement to depend directly on its context. Reviewing behavioral data, Dzhafarov et al. (2016) found no contextuality once direct context effects were separated out; Cervantes and Dzhafarov (2018) then gave the first clear demonstration of contextuality in human choice, and Basieva et al. (2019) showed that contextual systems can be found in tasks designed for the purpose, while earlier claims in judgments are explained by direct influences.

What it means for PBO. Incomplete preferences represented by a set of utilities correspond in PBO to a vector-valued Gaussian process with Pareto dominance (Section 14.5), close to Bayesian optimization with preference exploration, in which a person compares outcome vectors (Lin et al., 2022) (inference). A 50:50 answer is ambiguous between indifference, indecisiveness, and experimentation; an interface that offers "cannot decide" as a separate answer from "about the same", modeled separately from a tie, can tell them apart (inference; Section 20.4). Because most context effects are direct influences, a likelihood with covariates for order, choice set, and history can absorb them within one global utility; true contextuality appears only in specially designed tasks (inference).

Sources cited in Section 43.4 10
  1. Debreu (1964) Continuity Properties of Paretian Utility
  2. Hervés‐Beloso and del Valle‐Inclán Cruces (2019) Continuous preference orderings representable by utility functions
  3. Banerjee and Mitra (2018) On Wold’s approach to representation of preferences
  4. Evren and Ok (2011) On the multi-utility representation of preference relations
  5. Ok and Tserenjigmid (2022) Indifference, indecisiveness, experimentation, and stochastic choice
  6. Cettolin and Riedl (2019) Revealed preferences under uncertainty: Incomplete preferences and preferences for randomization
  7. Dzhafarov et al. (2016) Is there contextuality in behavioural and social systems?
  8. Cervantes and Dzhafarov (2018) Snow queen is evil and beautiful: Experimental evidence for probabilistic contextuality in human choices
  9. Basieva et al. (2019) True contextuality beats direct influences in human decision making
  10. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes

43.5 Information theory #

A common argument says that a pairwise comparison carries at most one bit, far less than a rating could, which makes comparisons inefficient. The bound is correct: the mutual information between a binary answer and the utility (Section 6.3) cannot exceed the entropy of the answer, at most 1 bit. It is not the binding constraint. Error rates for ordinal and cardinal estimation are identical up to constant factors (Shah et al., 2016). Per query, ordinal measurements have lower noise per sample and are usually faster to collect but usually carry less information, and a preprint by the same group (Shah et al., 2014) quantifies at which noise levels ordinal measurement is better. The one-bit argument concerns constants; the opposite argument, that regret bounds for PBO comparable to those of scalar BO show comparisons are not inefficient, concerns rates.

How much a comparison carries depends on the system and the person. Ghosal et al. (2023) showed that the "rationality coefficient" (the noise scale) should be fitted separately for each type of feedback, that overestimating human rationality severely harms the accuracy and regret of reward learning, and that when people are very suboptimal, comparisons are more informative than demonstrations. And efficient coding of value (Polanía et al., 2019), introduced in Section 39.4, allocates precision to the values a person expects to see.

What it means for PBO. The noise scale should be fitted per person and per type of feedback, not fixed at a low default (inference from Ghosal et al.). Under a probit or logistic link, a comparison's expected information approaches one bit only when the predicted win probability is near 0.5 and the noise is small (Exercise 43.2); adding ties or graded confidence raises the ceiling to log⁡23\log_2 3 bits for three answers but adds noise parameters (inference). As PBO narrows its candidates toward the optimum, efficient coding predicts that discrimination there may improve rather than hit a fixed floor, testable by fitting the slope of the psychometric function separately early and late in a session (inference). We found no direct measurement of how many bits a comparison carries for human design preferences.

Sources cited in Section 43.5 4
  1. Shah et al. (2016) Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
  2. Shah et al. (2014) When is it Better to Compare than to Score?
  3. Ghosal et al. (2023) The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
  4. Polanía et al. (2019) Efficient coding of subjective value

43.6 Game theory and symmetry breaking #

When two options are identical under the description the chooser uses, any choice between them must draw on information from outside that description, such as position, history, or salience (Schelling's focal points). A standard preference likelihood predicts 0.5 at equal utility and so treats every such choice as noise, which is misspecified if the symmetry breaking is systematic. In the network coordination experiments of Mäs and Nax (2016), 96% of decisions were myopic best responses, deviations were rarer when they cost more, and individuals differed. And choosing between options a person rated similarly, close to the symmetric case, itself changes preference: across 43 free-choice studies that excluded a known methodological artifact, Cohen's d=0.40d = 0.40 (95% confidence interval 0.32 to 0.49) (Enisman et al., 2021), as Section 37.2.2 describes.

Answers can also be strategic: every reasonable voting rule over three or more options can be manipulated by misreporting (the Gibbard-Satterthwaite theorem), and for rules on three alternatives that are far from a dictatorship and from having a range of only two alternatives, manipulation succeeds with non-negligible probability (Friedgut et al., 2011). A single binary query is strategy-proof, but misreports across a sequence remain possible.

What it means for PBO. Near indifference, real choices are pushed one way by position, salience, and defaults. Balancing the order of presentation and adding a position or default term to the likelihood, the prior weight ww of Equation (43.1), turns this from noise into a signal the model can estimate (inference). Choice-induced change means that a close comparison biases later answers toward the chosen option, a drift correlated with the optimizer's own queries (Section 42.2); asking some early close pairs again later, or adding a choice-history term to the likelihood, can detect it (inference). The payoff-sensitive deviations of Mäs and Nax support probit or logistic errors, whose rate falls as the utility difference grows, as the main mechanism of deviation, with a small lapse component (inference). We found no study of strategic misreporting in PBO.

Sources cited in Section 43.6 3
  1. Mäs and Nax (2016) A behavioral study of “noise” in coordination games
  2. Enisman et al. (2021) Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm
  3. Friedgut et al. (2011) A Quantitative Version of the Gibbard–Satterthwaite Theorem for Three Alternatives

43.7 Influence among users #

Complexity and network science also study how choices spread through a group, which matters once PBO has several users, a shared gallery, or a prior learned from earlier users. In the two "multiple worlds" experiments of Macy et al. (2019) (n=4,581n = 4{,}581), social influence made partisan divisions larger and less predictable: in parallel worlds, the same position became attached to opposite parties. Frey and van de Rijt (2021) found that when people choose one after another and see the running counts, majorities are more often wrong, and on hard tasks a wrong majority perpetuates itself instead of being corrected. And homophily (similar people connect) and contagion (connected people become similar) are generally confounded in observational social network studies (Shalizi and Thomas, 2011); the frequently heard version, that the two cannot be told apart even in principle, holds only for observational data.

What it means for PBO. For PBO with many users, gallery-style crowdsourcing, or transfer of earlier users' preference models to a new user (as in Meta-PO (Li et al., 2025a)), independent first judgments should be collected before any aggregate is shown, and a population prior trained on socially influenced data carries an arbitrary, path-dependent component (inference). Because homophily and influence cannot be separated in the logs of a shared gallery, the remedy is in the design: randomize which users see which earlier choices (inference from Shalizi and Thomas). For single-user PBO, the link is weak.

Sources cited in Section 43.7 4
  1. Macy et al. (2019) Opinion cascades and the unpredictability of partisan polarization
  2. Frey and van de Rijt (2021) Social Influence Undermines the Wisdom of the Crowd in Sequential Decision Making
  3. Shalizi and Thomas (2011) Homophily and Contagion Are Generically Confounded in Observational Social Network Studies
  4. Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences

43.8 Differential privacy #

Preference data are sensitive: the designs a person prefers can reveal their body, their health, or their tastes. A randomized algorithm is ε\varepsilon-differentially private if changing one person's data changes the probability of any output by at most a factor of eεe^{\varepsilon} (Dwork et al., 2006); smaller ε\varepsilon means stronger privacy. In local differential privacy, each answer is randomized before it leaves the user; in central differential privacy, a trusted curator holds the raw data and randomizes only what it releases. The classic local mechanism for a binary answer is randomized response (Warner, 1965): report the true answer with probability eε/(1+eε)e^{\varepsilon}/(1 + e^{\varepsilon}) and the opposite answer otherwise.

Since 2023, private learning from pairwise preferences has tight theoretical bounds. For a dd-dimensional Bradley-Terry reward estimated from nn comparisons under label privacy, Chowdhury et al. (2024) (AISTATS 2024) showed that the additional error is Θ((1/(eε−1))d/n)\Theta\big((1/(e^{\varepsilon} - 1)) \sqrt{d/n}\big) under local privacy and Θ(poly(d)/(εn))\Theta\big(\mathrm{poly}(d)/(\varepsilon n)\big) under central privacy, tight under local privacy and tight in nn and ε\varepsilon under central privacy. Matching bounds also exist, in preprints, for dueling bandits (Chapter 21) (Saha and Asi, 2024) and for offline reinforcement learning from human feedback (Wu et al., 2025b). On the attack side, PREMIA (Feng et al., 2025) found that models aligned with direct preference optimization (DPO) are more vulnerable to membership inference attacks on their preference data than models aligned with proximal policy optimization (Section 35.3). Private Bayesian optimization with scalar feedback exists (Kusner et al., 2015), but a search of arXiv abstracts for "preferential Bayesian optimization" together with "privacy" or "private" returned nothing.

What it means for PBO. The following are inferences. Randomized response on each comparison satisfies ε\varepsilon-local privacy, and the PBO likelihood can absorb it as label noise with a known flip probability q=1/(1+eε)q = 1/(1 + e^{\varepsilon}). The reported answer favors x\vx with probability q+(1−2q) pq + (1 - 2q)\,p, where pp is the unprivatized probability, so the slope of the answer probability with respect to the utility difference shrinks by 1−2q=tanh⁡(ε/2)1 - 2q = \tanh(\varepsilon/2). For a weak signal, with both probabilities near one half, the Fisher information of one answer shrinks by tanh⁡2(ε/2)\tanh^2(\varepsilon/2), and keeping the same information needs 1/tanh⁡2(ε/2)1/\tanh^2(\varepsilon/2) times as many queries: about 1.7 times at ε=2\varepsilon = 2, about 4.7 times at ε=1\varepsilon = 1, and about 17 times at ε=0.5\varepsilon = 0.5, consistent with the 1/(eε−1)1/(e^{\varepsilon} - 1) rate above. With only tens to hundreds of queries, that is expensive. In single-user PBO, what leaks a preference is the optimizer's output, the candidates it shows and the final design, so privacy of the outputs fits the threat better than privacy of the labels. Transferring preference models between users, as in Meta-PO, exposes earlier users' comparisons to membership inference of the kind PREMIA demonstrated, and there central privacy on the aggregated model is more practical than local privacy. The link is strong, but as of September 2026 we found no paper on private PBO.

Sources cited in Section 43.8 7
  1. Dwork et al. (2006) Calibrating Noise to Sensitivity in Private Data Analysis
  2. Warner (1965) Randomized Response: A Survey Technique for Eliminating Evasive Answer Bias
  3. Chowdhury et al. (2024) Differentially Private Reward Estimation with Preference Feedback
  4. Saha and Asi (2024) DP-Dueling: Learning from Preference Feedback without Compromising User Privacy
  5. Wu et al. (2025b) Offline and Online KL-Regularized RLHF under Differential Privacy
  6. Feng et al. (2025) Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment
  7. Kusner et al. (2015) Differentially Private Bayesian Optimization

43.9 Common claims, checked #

Table 43.1 collects the formal claims checked in this chapter that bear on how a PBO system is built.

Table 43.1 Claims from the natural and formal sciences, and what the sources say.
Claim What the sources say Section
The PBO likelihood is a softmax, a Boltzmann distribution only with a logistic link; Chu and Ghahramani and BoTorch's default use the probit Section 43.2
Luce's axiom and Thurstone's Case V are equivalent under a logistic distribution (Yellott 1977) Luce's form follows from Gumbel errors, whose differences are logistic; Case V assumes normal errors Section 43.2
Slime molds choose intransitively their ranking was linear and transitive Section 43.1
The lexicographic order on R2\R^2 has no continuous representation it has no real-valued representation at all Section 43.4.1
One bit per comparison makes comparisons inefficient the bound holds, but ordinal and cardinal error rates differ only by constants Section 43.5
Homophily and influence are indistinguishable in principle confounded in observational data Section 43.7

43.10 Settled, contested, missing #

Research status Settled, contested, missing

Settled. Estimation error from comparisons depends on the comparison graph's Laplacian spectrum, and the relative error scales with the square root of the effective resistance, with matching lower bounds up to logarithmic factors. Ordinal and cardinal error rates differ only by constants, although a binary answer carries at most one bit. Comparison data split orthogonally into a part one utility explains and cyclic parts. The PBO likelihood is a Boltzmann distribution under the logistic link, not under the probit. Choosing between similar options changes preferences (Cohen's d=0.40d = 0.40 across 43 studies). Private learning from pairwise comparisons has tight rates (some of them still in preprints).

Contested. How intransitive real preferences are: animal evidence mostly shows context or state dependence on a transitive core, and decoy effects in bees are found in one design and not another. Whether contextuality, beyond direct context effects, exists in ordinary choices. Whether utility scales adapt fast enough to change within a single session.

Missing. Spectral bounds of comparison graphs applied to Gaussian process PBO. A Hodge analysis of real PBO logs against its chance level. A fit of prior-dependent logit models to PBO's pairwise data. A measurement of the bits per comparison for human design preferences. A test of whether discrimination sharpens late in a session. A study of strategic misreporting in PBO. A private PBO method.

43.11 Exercises #

Exercise 43.1

In the Incumbent vs challenger design of Figure 43.1, every challenger is compared once with the incumbent. Use the rules for resistors to find the standard deviation of the utility difference between two challengers, and between a challenger and the incumbent. What happens to both when every comparison is asked twice?

Solution

A challenger and the incumbent are joined by one one-ohm resistor: resistance 1, standard deviation 1. Two challengers are joined through the incumbent by two resistors in series: resistance 2, standard deviation 2≈1.41\sqrt{2} \approx 1.41. Asking every comparison twice halves every resistance, giving 1/2≈0.71\sqrt{1/2} \approx 0.71 (the value the figure shows at 22 comparisons) and 1.

Exercise 43.2

Under a probit model, the person prefers x\vx with probability p=Φ(d/σ)p = \Phi(d/\sigma), where dd is the utility difference and σ\sigma the noise. The model is uncertain about dd: suppose d=+δd = +\delta or −δ-\delta with equal probability. Show that the mutual information between the answer and dd is 1−H2(Φ(δ/σ))1 - H_2(\Phi(\delta/\sigma)) bits, where H2H_2 is the binary entropy. When does it approach one bit, and what is it when δ=σ\delta = \sigma?

Solution

The mutual information is the entropy of the answer minus its expected conditional entropy, I=H(answer)−Ed[H(answer∣d)]I = H(\text{answer}) - \E_d[H(\text{answer} \mid d)] (Section 6.3). By symmetry the answer has entropy 1 bit overall, and given either sign of dd it has entropy H2(Φ(δ/σ))H_2(\Phi(\delta/\sigma)), so I=1−H2(Φ(δ/σ))I = 1 - H_2(\Phi(\delta/\sigma)). It approaches one bit when δ/σ\delta/\sigma is large: the model is unsure of the sign but the person answers almost deterministically. At δ=σ\delta = \sigma, Φ(1)≈0.841\Phi(1) \approx 0.841 and H2(0.841)≈0.63H_2(0.841) \approx 0.63, so the answer carries about 0.37 bits.

Exercise 43.3

A study plans 60 comparisons per participant and wants each answer protected by randomized response with ε=1\varepsilon = 1. Using the calculation in Section 43.8, roughly how many comparisons would carry the same information as the 60 unprotected answers? What does this suggest about where privacy should be applied in a single-user session?

Solution

The information per weak-signal answer shrinks by tanh⁡2(1/2)≈0.4622≈0.21\tanh^2(1/2) \approx 0.462^2 \approx 0.21, so matching 60 unprotected answers needs about 60/0.21≈28060 / 0.21 \approx 280 protected ones, more than four times the budget. Because a single user is both the data subject and the person answering, the more economical protection is on what the system releases (the candidates shown to others, the stored model, the final design), not on each answer.

Further reading #

References

  1. Armand, M., Herrnberger, L., Jung, C., and Czaczkes, T. J. (2026). No evidence of a decoy effect in bees: Rewardless flowers do not increase bumblebees' preference for neighbouring flowers. Ecological Entomology. doi:10.1111/een.70092. Cited in §43.1
  2. Banerjee, K., and Mitra, T. (2018). On Wold’s approach to representation of preferences. Journal of Mathematical Economics. doi:10.1016/j.jmateco.2018.08.007. Cited in §43.4
  3. Basieva, I., Cervantes, V. H., Dzhafarov, E. N., and Khrennikov, A. (2019). True contextuality beats direct influences in human decision making. Journal of Experimental Psychology: General. Cited in §43.4
  4. Cervantes, V. H., and Dzhafarov, E. N. (2018). Snow queen is evil and beautiful: Experimental evidence for probabilistic contextuality in human choices. Decision. Cited in §43.4
  5. Cettolin, E., and Riedl, A. (2019). Revealed preferences under uncertainty: Incomplete preferences and preferences for randomization. Journal of Economic Theory. Cited in §43.4
  6. Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology. Cited in §43.1
  7. Chau, S. L., González, J., and Sejdinovic, D. (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. Cited in §43.3
  8. Chowdhury, S. R., Zhou, X., and Natarajan, N. (2024). Differentially Private Reward Estimation with Preference Feedback. International Conference on Artificial Intelligence and Statistics. Cited in §43.8
  9. Debreu, G. (1964). Continuity Properties of Paretian Utility. International Economic Review. Cited in §43.4
  10. Dwork, C., McSherry, F., Nissim, K., and Smith, A. (2006). Calibrating Noise to Sensitivity in Private Data Analysis. Theory of Cryptography (TCC 2006), Lecture Notes in Computer Science. Cited in §43.8
  11. Dzhafarov, E. N., Zhang, R., and Kujala, J. (2016). Is there contextuality in behavioural and social systems? Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences. Cited in §43.4
  12. Enisman, M., Shpitzer, H., and Kleiman, T. (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Cited in §43.6
  13. Evren, Ö., and Ok, E. A. (2011). On the multi-utility representation of preference relations. Journal of Mathematical Economics. Cited in §43.4
  14. Feng, Q., Kasa, S. R., Kasa, S. K., Yun, H., Teo, C. H., and Bodapati, S. B. (2025). Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment. International Conference on Artificial Intelligence and Statistics. Cited in §43.8
  15. Fiedler, M. (1973). Algebraic Connectivity of Graphs. Czechoslovak Mathematical Journal. Cited in §43.3
  16. Frey, V., and van de Rijt, A. (2021). Social Influence Undermines the Wisdom of the Crowd in Sequential Decision Making. Management Science. Cited in §43.7
  17. Friedgut, E., Kalai, G., Keller, N., and Nisan, N. (2011). A Quantitative Version of the Gibbard–Satterthwaite Theorem for Three Alternatives. SIAM Journal on Computing. Cited in §43.6
  18. Ghosal, G. R., Zurek, M., Brown, D. S., and Dragan, A. D. (2023). The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types. AAAI. Cited in §43.5
  19. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §43.2
  20. Harhen, N. C., and Bornstein, A. M. (2023). Overharvesting in human patch foraging reflects rational structure learning and adaptive planning. Proceedings of the National Academy of Sciences. Cited in §43.1
  21. Heckel, R., Shah, N. B., Ramchandran, K., and Wainwright, M. J. (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. Cited in §43.3
  22. Hemingway, C. T., DeVore, J. E., and Muth, F. (2024). Economic foraging in a floral marketplace: asymmetrically dominated decoy effects in bumblebees. Proceedings of the Royal Society B: Biological Sciences. Cited in §43.1
  23. Hendrickx, J. M., Olshevsky, A., and Saligrama, V. (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. Cited in §43.3
  24. Hervés‐Beloso, C., and del Valle‐Inclán Cruces, H. (2019). Continuous preference orderings representable by utility functions. Journal of Economic Surveys. Cited in §43.4
  25. Jiang, X., Lim, L.-H., Yao, Y., and Ye, Y. (2011). Statistical ranking and combinatorial Hodge theory. Mathematical Programming. Cited in §43.3
  26. Kusner, M. J., Gardner, J. R., Garnett, R., and Weinberger, K. Q. (2015). Differentially Private Bayesian Optimization. International Conference on Machine Learning. Cited in §43.8
  27. Latty, T., and Beekman, M. (2011). Irrational decision-making in an amoeboid organism: transitivity and context-dependent preferences. Proceedings of the Royal Society B: Biological Sciences. Cited in §43.1
  28. Li, Z., Liao, Y.-C., and Holz, C. (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Cited in §43.7
  29. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §43.4
  30. Lindig-León, C., Kaur, N., and Braun, D. A. (2022). From Bayes-optimal to heuristic decision-making in a two-alternative forced choice task with an information-theoretic bounded rationality model. Frontiers in Neuroscience. Cited in §43.2
  31. Luce, R. D. (1959). Individual Choice Behavior: A Theoretical Analysis. Wiley. Cited in §43.2
  32. Macy, M., Deri, S., Ruch, A., and Tong, N. (2019). Opinion cascades and the unpredictability of partisan polarization. Science Advances. Cited in §43.7
  33. Mäs, M., and Nax, H. H. (2016). A behavioral study of “noise” in coordination games. Journal of Economic Theory. Cited in §43.6
  34. Matějka, F., and McKay, A. (2015). Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model. American Economic Review. Cited in §43.2
  35. Meta Platforms, Inc. (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. software Cited in §43.2
  36. Netzer, N. (2009). Evolution of Time Preferences and Attitudes toward Risk. American Economic Review. Cited in §43.1
  37. Ok, E. A., and Tserenjigmid, G. (2022). Indifference, indecisiveness, experimentation, and stochastic choice. Theoretical Economics. Cited in §43.4
  38. Polanía, R., Woodford, M., and Ruff, C. C. (2019). Efficient coding of subjective value. Nature Neuroscience. Cited in §43.5
  39. Saha, A., and Asi, H. (2024). DP-Dueling: Learning from Preference Feedback without Compromising User Privacy. arXiv. preprint Cited in §43.8
  40. Schuck-Paim, C., Pompilio, L., and Kacelnik, A. (2004). State-Dependent Decisions Cause Apparent Violations of Rationality in Animal Choice. PLoS Biology. Cited in §43.1
  41. Shafir, S. (1994). Intransitivity of preferences in honey bees: support for 'comparative' evaluation of foraging options. Animal Behaviour. Cited in §43.1
  42. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. (2014). When is it Better to Compare than to Score? arXiv. preprint Cited in §43.5
  43. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. J. (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Cited in §43.3 §43.5
  44. Shalizi, C. R., and Thomas, A. C. (2011). Homophily and Contagion Are Generically Confounded in Observational Social Network Studies. Sociological Methods & Research. Cited in §43.7
  45. Strang, A., Abbott, K. C., and Thomas, P. J. (2022). The Network HHD: Quantifying Cyclic Competition in Trait-Performance Models of Tournaments. SIAM Review. Cited in §43.3
  46. Warner, S. L. (1965). Randomized Response: A Survey Technique for Eliminating Evasive Answer Bias. Journal of the American Statistical Association. Cited in §43.8
  47. Wu, Y., Thareja, R., Vepakomma, P., and Orabona, F. (2025b). Offline and Online KL-Regularized RLHF under Differential Privacy. arXiv. preprint Cited in §43.8
  48. Yellott, J. J. I. (1977). The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution. Journal of Mathematical Psychology. Cited in §43.2