Designing the Question
Chapter 19 built preferential Bayesian optimization (PBO) around one question: of these two options, which do you prefer? The model chose the two options, the person answered with one bit, and the loop repeated. That question is a design decision, and it is not the only one available. A system can show four options and ask for the best, or ask for an order. It can hand the person a slider and let them search a whole line of designs in one gesture. It can let them say that two options look about the same, that they are not sure, or that the experiment failed.
Each of these choices changes three things at once: how much effort an answer costs the person, how much the answer can tell the model, and what kind of noise the answer carries. The likelihood, the function that says how probable each answer is given the utility, has to change with it, because the likelihood is the model's description of a person using a particular interface. This chapter works through the main alternatives to the pair, each with its likelihood, and ends with the argument that gives the chapter its shape: the interface is part of the model.
20.1 Pairs, sets, and rankings #
A pair is the smallest question that reveals a preference, and that is both its strength and its limit. It is easy to answer, but its answer carries at most one bit, and in a design space of several dimensions one bit is a small step. The obvious extension is to show more options at once. Showing options and asking for the best one can carry up to bits; asking for a complete order of options can carry up to bits, which is about 4.6 bits for four options. These are upper bounds, reached only when every answer is equally likely in advance (Section 16.6), but they show what is at stake.
20.1.1 Choosing one from a set #
The model for a choice among several options goes back to Luce (1959). Let be the set of options shown and the latent utility. The probability that the person picks is
where , the scale of the logistic link (Equation (16.4)), often called a temperature, sets how noisy the choices are: a small makes the person pick the best option almost always, a large makes the choice nearly uniform. A machine-learning reader will recognize the right-hand side as the softmax function. With two options it reduces to , the logistic link of the Bradley-Terry model (Bradley and Terry, 1952) that Section 16.4 introduces.
The formula has a random-utility reading (Section 16.5). Suppose the person perceives each option's utility with independent noise drawn from a Gumbel distribution, the distribution that describes the largest of many random values, and picks the option whose perceived utility is highest. Then the probability that wins is exactly Equation (20.1) (McFadden, 1974). The pairwise probit model of Chapter 18 makes the same assumption with Gaussian noise instead; for two options the two links differ only slightly in shape, but only the Gumbel version gives a closed form for larger sets.
Equation (20.1) has a consequence that matters for interface design. It inherits the independence of irrelevant alternatives that follows from Luce's choice axiom (Section 16.4): the ratio of the probabilities of choosing and is , whatever else is in the set. Real choices sometimes violate it, as in the decoy or attraction effect (Huber et al., 1982). A pair has no third option, so it cannot show this effect; a set of four can. The effect appears mostly when each attribute of the options is shown as a number, and usually not when the attributes are perceived directly (Frederick et al., 2014). That is reassuring for visual design but not a guarantee (inference). Section 37.2 reviews this evidence.
With Equation (20.1) as the likelihood, the rest of the machinery is unchanged. The Laplace approximation of Section 18.2 needs the gradient and the curvature of the log-likelihood. For one choice with probabilities over the options in , the gradient with respect to is , where is 1 for the chosen option and 0 otherwise, and the negative Hessian is on those options: the covariance matrix of a one-hot vector drawn with probabilities . It is positive semidefinite, as the Laplace approximation requires, and for two options it is the familiar pairwise term.
20.1.2 Does a larger set help? #
The evidence on larger sets points in two directions. Siivola et al. (2021) derived a likelihood for any parallel feedback on two or more points and argued that the batch winner, the best option of a set, is the most useful form, because full rankings of large batches are laborious for people and sometimes impossible, as in A/B testing. On six benchmark functions with batches of four, they found that the differences between acquisition functions were larger than the differences between feedback types; across batch sizes from 2 to 6 they saw no clear difference. Their experiments used at most four dimensions. The qEUBO experiments (Section 19.4), with an acquisition function designed for sets, found the opposite: queries of four options reached a given simple regret (the gap between the best utility and that of the recommended option, Section 13.1) with fewer queries than pairs, while the further gain from four to six was smaller and less stable; the authors note the contrast with Siivola et al. (Astudillo et al., 2023). The difference may lie in the acquisition functions or the noise levels, but no paper has separated these factors (inference).
Theory for finite sets of options gives a sharper answer, and it depends on what the person reports. Under the Plackett-Luce model introduced below, and for a finite set of options, consider the sample complexity of finding a near-best option: the number of queries needed to find one with high probability. Learning from the winners of -option sets has the same sample complexity, up to constant factors, as learning from pairs: it does not improve with . Reporting the top of each set instead reduces it by a factor of (Saha and Gopalan, 2019b). A larger set helps only when the person tells more than which option won. For linear utilities, a 2025 result shows that larger subsets provably help under the Plackett-Luce model (Lee et al., 2025a); as of September 2026 this has not been carried over to the Gaussian process models of this book. Chapter 21 returns to these results.
20.1.3 Rankings #
A ranking can be read as a sequence of choices: first the best of the set, then the best of what remains, and so on. Under the choice axiom, each later choice is again Equation (20.1) on the remaining options. That gives the probability of a whole ranking.
Let order the shown options from best to worst, so is ranked first.
- By the chain rule of probability (Section 2.4), the probability of the ranking is the probability of the first place, times the probability of the second place given the first, and so on: a product over places of the probability that takes place given the places before it.
- Given the first places, the -th place is a choice of the best option among the ones not yet placed, .
- Assume, in the spirit of Luce's choice axiom, that the options already placed do not affect that choice. Then it follows Equation (20.1) with the set .
- Substituting gives Equation (20.2). The last factor, , is a choice from a set of one and equals 1.
This is the Plackett-Luce model (Plackett, 1975; Luce, 1959). A top- ranking, in which the person orders only the best options, keeps the first factors. The Gumbel reading carries over: if every option's perceived utility has independent Gumbel noise and the person sorts by perceived utility, the ranking follows Equation (20.2).
Nguyen et al. (2021) built a Gaussian process surrogate inspired by the multinomial logit and its ranking extension, which can be read as Gaussian process regression with independent Gumbel noise; it handles top- rankings and ties, is trained by variational inference, and was evaluated on synthetic functions, CIFAR-10, and the SUSHI preference data. The tutorial of Benavoli and Azzimonti (2026a) describes it as a Gaussian process generalization of the Plackett-Luce model. Benavoli et al. (2023) went further and let the person choose a subset (of five options A to E, say, the acceptable ones are A, B, and C), modeling the choice with several latent utilities and choosing their number by how well the model predicts held-out choices (cross-validation, Section 22.1.2); the abstract reports simulations only. As of September 2026 we found no newer Gaussian process PBO paper with a Plackett-Luce or multinomial ranking likelihood.
Showing more options raises what an answer can carry, but only if the person reports more than the winner. A set of four with only its winner reported is, in the worst case, no better than a pair; a ranking of the same four can be.
Sources cited in Section 20.1 13
- Luce (1959) Individual Choice Behavior: A Theoretical Analysis
- Bradley and Terry (1952) Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
- McFadden (1974) Conditional Logit Analysis of Qualitative Choice Behavior
- Huber et al. (1982) Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis
- Frederick et al. (2014) The Limits of Attraction
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Saha and Gopalan (2019b) PAC Battling Bandits in the Plackett-Luce Model
- Lee et al. (2025a) Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
- Plackett (1975) The Analysis of Permutations
- Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
- Benavoli and Azzimonti (2026a) A tutorial on learning from preferences and choices with Gaussian Processes
- Benavoli et al. (2023) Learning Choice Functions with Gaussian Processes
20.2 Searching along a line #
Sets and rankings still compare options the system chose. In a design space of six or ten parameters, the system's few options are a sparse sample, and most of the person's knowledge about what would look better goes unused. The opposite extreme, handing the person every parameter, does not work either: exploring a high-dimensional space directly is difficult even for designers (Koyama and Igarashi, 2018), and absolute scores fail for the reason given in Section 16.1.1: they require a familiarity with the whole design space that a newcomer does not have.
Koyama et al. (2017) found a middle ground: give the person a single slider. The slider maps to a line segment in the design space, the preview updates as the person drags, and the person stops where the design looks best. One answer is a one-dimensional optimization performed by the person, over a continuum of designs, with no familiarity with the parameters required. The method is called sequential line search.
20.2.1 Which line #
The system chooses the segment. After answers, with a Gaussian process posterior over the utility, let be the observed design with the highest posterior mean and the design that maximizes expected improvement (Section 12.3). The next slider runs between them:
One end is the best design so far, so the person can always keep it. The other is the most promising place to look, so the slider spans the trade-off between exploiting and exploring that every acquisition function balances. The first slider, before any data, connects two random points (Koyama and Igarashi, 2018).
The answer is the chosen point on the segment. Koyama et al. record it as a choice from a set of three: the chosen design beats both ends, , with the likelihood of Equation (20.1) on those three designs (Koyama et al., 2020). The record throws information away. The person preferred the chosen point to every other point on the slider, a continuum of comparisons, not just to the two ends. The authors note that more points could be added to the losing side at extra computational cost. Mikkola et al. (2020) take the continuum seriously: they treat the answer as uncountably many pairwise comparisons, write its likelihood as a limit of products over finer and finer partitions of the line, and approximate it with a finite set of sampled comparisons.
Input: a design space , kernel , choice scale , budget sliders.
- Set the first slider between two random designs.
- Show the slider with a live preview. Record the position the person chooses, .
- Add the choice to the data and fit the utility posterior with the likelihood Equation (20.1) and the Laplace approximation.
- Find , the observed design with the highest posterior mean, and , the design with the highest expected improvement over it.
- Set the next slider by Equation (20.3) and return to step 2 until the budget is spent. Recommend .
20.2.2 Try it #
The figure below runs Algorithm 20.1 with you on the slider. The design is a small poster landscape with two parameters, warmth and vividness. Move the slider, or drag along the strip of thumbnails, until the picture looks best to you, then press Choose this one. The map on the right shows the model's posterior mean over the two parameters (darker is better), the sliders you have used, the designs you chose, and the next slider in orange, running from the best design so far (circle) to the point of highest expected improvement (diamond).
Some things to try. Answer five or six sliders honestly and watch the sliders on the map: they get shorter and cluster as the model becomes sure where your favorite region is, and the far end jumps to a new region when the uncertainty elsewhere makes expected improvement large. Choose an end of the slider without moving it, and the record becomes an ordinary pair, the best design against the expected-improvement design. In the simulated mode, reveal the hidden favorite and press Simulate five a few times: with the default noise, the star usually sits close to the hidden favorite by the tenth slider. Raise the noise to 0.2 and the chosen points scatter along each slider; the model, which assumes a fixed , reads the scatter as weak preferences and converges more slowly.
The map makes the method look easy, because in two dimensions a dozen lines already pass close to most of the space. The method was built for more. In six dimensions, as in the photo task, a dozen lines pass close to only a small part of the space, so the choice of line does the work: one end anchors the search at the best design so far, and expected improvement points the other end at the region where a better design is most likely. The Sequential Gallery simulations below, run in 5 to 20 dimensions, show how much the choice of subspace matters there.
20.2.3 What the crowd did #
Sequential line search was designed for crowdsourcing, in which many paid workers each perform a small task (a microtask) and the system combines their answers. Koyama and Igarashi (2018) give the procedure in detail. Every result used 15 iterations. In each iteration the system posted seven slider microtasks and moved on once at least five answers had arrived, using the median of the returned slider positions as the choice. A microtask paid 0.05 USD, so a result cost 5.25 USD, and the photo examples took about 68 minutes on average.
For photo color enhancement with six parameters, crowd workers were then asked which of four versions of each photo looked best: the original, the crowd-optimized result, and the automatic enhancements of Adobe Photoshop and Lightroom. Across three photos the crowd-optimized versions received 32, 26, and 29 votes, against 0 to 3 for each alternative. Three runs started from different initial conditions produced similar enhancements, and the differences between them shrank rapidly within the first four or five iterations (Koyama and Igarashi, 2018). Chapter 25 works through a photo enhancement problem of the same kind, with pairs and with a slider.
Two features of this setup recur in the rest of the chapter. Averaging over many workers assumes, in the authors' words, that "a common 'general' preference exists that is shared among crowds" (Section 20.5). And a slider answer has its own kind of noise: how precisely a person can position a slider, and how finely they can see differences between nearby designs, are properties of the person and the widget, not of the utility. The papers on slider and projection queries fold it into the same single noise term as everything else (inference).
Sources cited in Section 20.2 4
- Koyama and Igarashi (2018) Computational Design with Crowds
- Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
20.3 Galleries and projections #
A line is one way to let the person search a subspace. Two other designs generalize it: a plane of designs shown as a grid, and a projection along any direction the system chooses.
20.3.1 Sequential Gallery #
Koyama et al. (2020) replaced the slider with a two-dimensional plane and the preview with a zoomable grid: a gallery of designs sampled from the plane, 5 by 5 for the photo task. The person clicks the best design; the grid zooms in around it by a factor of two; after a fixed number of clicks (four in their implementation) the selection is the answer. The method is called sequential plane search, and the whole interactive framework Sequential Gallery.
The plane is chosen like the line. It is centered at the current best design , has the expected-improvement point as one of its vertices, and is otherwise placed to maximize a new acquisition function, the average expected improvement over the plane, approximated on a 5 by 5 lattice of points. The answer is recorded as a choice over five representative points: the chosen design beats the center and the four vertices.
In simulations with 50 trials per method, on test functions of 5 to 20 dimensions, plane search outperformed line search on every function at every iteration count, and the acquisition-based plane outperformed a random plane after the first few iterations. A preliminary study with five students and one researcher (one participant described themselves as an expert, the other five as novices) enhanced photos: participants pressed a button to say they were satisfied after 5.36 iterations on average (standard deviation 2.69), and one plane subtask took 14.8 seconds on average. The statement "I could get inspiration for possible enhancement from the grid view" scored a mean of 6.50 (standard deviation 0.548) on a seven-point scale. The authors list the limitations themselves: the grid suits only designs that can be recognized at a glance, discrete parameters are not handled, and Bayesian optimization is known to perform poorly beyond about 20 dimensions (Koyama et al., 2020).
20.3.2 Projective preferential queries #
Mikkola et al. (2020) asked the person for the best position along a projection. A query is a direction and a reference design ; the person reports the best scalar for the design , with the coordinates that leaves at zero held at the values in . When is a coordinate direction, this is "set this one knob to its best value"; in their user experiment, materials scientists set one coordinate at a time of the position and orientation of a molecule above a surface.
Their likelihood is the continuum version of the slider record. They assume Thurstone's Gaussian noise, modeled as a white-noise process along the projection, write the probability that the reported point beats every other point on the projection as a product integral, and approximate it with sampled pseudo-comparisons inside the Laplace approximation. With a budget of 100 queries on four test functions of 2, 6, 10, and 20 dimensions, a preferential coordinate descent strategy was best on three of the four, a projective version of expected improvement was best on the 20-dimensional function, and all projective variants clearly outperformed all pairwise variants. To show how little a pair carries, they trained the dueling model of González et al. (2017) (Section 19.2) on 2000 random duels of the two-dimensional six-hump camel function; the point it found had value 0.1052 against the global minimum of −1.0316, and optimizing its soft-Copeland score took 41 minutes, while the projective version with random queries reached that accuracy within its first queries.
The paper also states two cautions that belong in this chapter. The more nonzero coordinates a direction has, "the greater the 'cognitive burden' to a human user", so the best query for a person is not the best query for a perfect oracle. And a person may be unable to state a best value along some direction at all; the authors suggest allowing the answer "I do not know" and leave it for future research (Mikkola et al., 2020).
20.3.3 Lines for the algorithm, not the person #
A line can also restrict the algorithm rather than the person. LineCoSpar, for exoskeleton gait tuning, takes at each iteration a random line through the design with the highest posterior mean and runs Thompson sampling (Section 12.5) over that line and the designs already visited; the person still answers pairwise comparisons of gaits. Tested with six able-bodied participants tuning six gait parameters, it was introduced because running the group's earlier method, CoSpar (Tucker et al., 2020b), on a six-dimensional space was infeasible (Tucker et al., 2020a). The idea is the same as in Section 20.2, a one-dimensional subspace chosen to be promising, used for the computation instead of the interaction. Chapter 30 follows these subspace methods further.
Table 20.1 collects the forms so far. The column of bits is the logarithm of the number of possible answers, a ceiling that a noisy answer never reaches; it shows how the forms differ in what they could carry, not what they do carry.
| Form | The person | Recorded as | At most | Example |
|---|---|---|---|---|
| Pair | picks one of two | 1 bit | Brochu et al. (2007); Chapter 19 | |
| Best of | picks one of | Equation (20.1) | bits | Siivola et al. (2021); Astudillo et al. (2023) |
| Top- of | orders the best | first factors of Equation (20.2) | bits | Nguyen et al. (2021) |
| Slider | drags to the best point on a segment | chosen beats both ends, or a continuum | of the slider's resolution | Koyama et al. (2017) |
| Gallery | clicks the best of a 5 by 5 grid, four times | chosen beats center and vertices | bits | Koyama et al. (2020) |
| Projection | sets a direction to its best value | a continuum of comparisons | resolution-limited | Mikkola et al. (2020) |
| With "about the same" | picks a, b, or neither | Equation (20.4) | bits | Bıyık et al. (2019) |
| Graded | picks one of ordered levels, from "a, clearly" to "b, clearly" | Equation (20.6) | bits | Li et al. (2021); Wu et al. (2025a) |
20.3.4 What the evidence says about forms #
Read across papers, the evidence has a consistent direction and a consistent gap. Forms that let one human action carry more (projections, planes, queries of four) beat pairs in the papers that introduced them, and the only direct comparison between batch winners and full rankings found little difference (Siivola et al., 2021). But almost every comparison is a simulation, and each paper uses its own acquisition function, budget, and simulated noise. The first same-task comparisons with people appeared in 2026: in GimmBO, with 12 participants (all with computer science or machine learning backgrounds) and 20 iterations each, ranking beat a slider on image similarity and success rate; the slider and gallery baselines ended with redundant adapters switched on, that is, adapters outside the set that had produced the target image (an adapter is a small fine-tuned add-on to the image model, and the task was to tune the weights with which adapters are merged); and users of Sequential Gallery reported getting stuck in local minima. Ranking took longer per step, 50.5 seconds against 34.7 for the slider and 10.6 for the gallery (Liu et al., 2026b). As of September 2026 we found no controlled study with people that compares pairs, batch winners, rankings, and sliders under the same acquisition function and budget, and no acquisition function derived for slider, plane, or projection queries beyond the adapted expected improvement and random subspaces described here. Section 32.4 reports the human-factors evidence in detail.
Sources cited in Section 20.3 14
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- González et al. (2017) Preferential Bayesian Optimization
- Tucker et al. (2020b) Preference-Based Learning for Exoskeleton Gait Optimization
- Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
- Brochu et al. (2007) Active Preference Learning with Discrete Choice Data
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
- Koyama et al. (2017) Sequential line search for efficient visual design optimization by crowds
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
- Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
- Wu et al. (2025a) Mixed Likelihood Variational Gaussian Processes
- Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
20.4 Ties, indifference, and confidence #
Every model so far forces an answer. A person who sees two designs that look the same to them must still pick one, and what they pick is close to a coin flip. The model cannot tell a coin flip from a weak preference, so it reads the flip as evidence that one option is slightly better. This section looks at answers that say more than "a" or "b": about the same, not sure, how sure, and the experiment failed.
20.4.1 About the same #
The simplest extension gives the person a third button. To model it, start from Thurstone's picture (Section 16.3): the person perceives the utility difference with Gaussian noise of standard deviation . Add an indifference threshold , a just-noticeable difference in the sense of psychophysics (Section 16.2): differences smaller than are not reported as preferences. The three answers then have probabilities
With the middle answer never occurs and the model is the probit pair of Chapter 18. The tutorial of Benavoli and Azzimonti (2026a) lists a just-noticeable-difference likelihood of this kind among its nine models, with indistinguishability statements of the form , tracing the idea to Luce's 1956 notion of a discrimination threshold. Erarslan et al. (2025) use such a threshold in an extended Thurstone model and report clear gains when 10% to 20% of comparisons are indifferent; theirs is a 2025 preprint.
The logistic version is due to Bıyık et al. (2019), who added an "About Equal" option to active reward learning with a minimum perceivable difference :
It reduces to the Bradley-Terry model at . Their reward was a linear function of trajectory features rather than a Gaussian process, and in their simulations the weak-preference queries consistently reduced the number of wrong answers. The two treatments in use, a threshold in a Thurstone or logistic model and ties in a multinomial logit as in Nguyen et al. (2021), are the ones the field has settled on (inference).
How much is the third button worth? Information theory answers this cleanly (Section 6.3). Compare two interfaces shown to the same person, whose answers follow Equation (20.4). With three buttons, they report what they perceive. With two, they report "a" or "b" when they perceive a preference and flip a coin when they do not.
Let be the unknown utility difference, the three-button answer, and the two-button answer.
- is computed from alone: when is "a" or "b", and is a fair coin when is "same". The coin does not depend on .
- So is a chain in which depends on only through .
- The data-processing inequality says that processing an observation can only lose information about its cause: for such a chain, , where is mutual information (see Cover and Thomas, 2006, ch. 2).
- The inequality is usually strict, because "same" itself carries information (that is probably small) and the coin erases it.
The figure shows the size of the loss. The top panel plots the three answer probabilities of Equation (20.4) against the utility difference, with the forced-choice probability of "a" dashed. The bottom panel plots how many bits one answer carries about , for each interface, when the model's current belief about is Gaussian with mean and a given spread.
Some things to try. Set the threshold to zero: the two curves in the bottom panel coincide, because nobody is ever indifferent. Raise it to 1.5 with the model's guess at , a close pair: the forced choice keeps only about a third of what the three-button answer carries. Move to 3, a pair the model already thinks is lopsided: both interfaces carry less than they do at (0.33 bits against 0.57 with three buttons, 0.14 against 0.19 with two), because the answer is more predictable. Shrink the uncertainty about : the bits fall toward zero for both, which is why acquisition functions avoid asking about pairs the model is already sure of.
Two cautions keep the figure honest. A model that has no "same" outcome, fitted to forced-choice data from a person who is often indifferent, reads the coin flips as noise: if the noise level is fitted, it grows; if it is fixed, the estimated utility differences shrink and the whole estimate flattens (inference). And a third button may change behavior as well as recording it: people may press it to avoid an effortful judgment, a possibility the threshold model does not describe (inference).
20.4.2 Not sure, and how sure #
"About the same" says the difference is small. "I do not know" says something different: that the person cannot judge, perhaps because the options differ in ways they cannot weigh. Mikkola et al. proposed that answer and left it for future work (Section 20.3.2), and an IUI 2023 study let participants rank four candidates while putting candidates they could not judge into a separate "don't know" area, to express incomplete preferences (Ou et al., 2023). As of September 2026 we found no Gaussian process PBO paper whose model has an abstain or skip outcome distinct from a tie.
Confidence ratings go the other way and say more. A graded answer such as "a, clearly" or "a, slightly" is an ordinal answer, and the threshold model extends to it directly: place cut points on the perceived difference and report level when it falls between and ,
Equation (20.4) is the case of three levels with cut points and . A model can also add a lapse rate , the probability that an answer is a random slip, so that each level has probability times the expression above; a single careless click then cannot drag the posterior far.
Three studies show the range of what has been done. ROIAL combined pairwise preferences with ordinal labels (very bad, bad, neutral, good) for exoskeleton gaits, arguing that an -level ordinal query yields at most bits against one bit for a preference; it was tested with 3 participants tuning 4 gait parameters (Li et al., 2021). Wu et al. (2025a), a preprint, combined the probit preference likelihood with a Likert confidence likelihood (a Likert item is a rating on a fixed scale of labeled levels) with learnable cut points and a lapse rate, and on human comparisons of robot gaits report consistently lower Brier scores and higher F1 scores, two measures of predictive accuracy. In a study of vibrotactile feedback, 13 participants made 40 rounds of pairwise comparisons with a five-level confidence rating that set the noise scale of each comparison, and the learned model reached a held-out accuracy of 92.3% (range 85% to 100%) (Zhang et al., 2026b). Response times, which can carry similar information without asking anything extra, are covered with the rest of these extensions in Section 27.2.
20.4.3 It crashed #
Some evaluations produce no design to judge. A controller makes a robot fall, a simulation diverges, a recipe is inedible. Forcing such an outcome into a comparison ("the crash is worse than anything") puts it on the utility scale where it does not belong. The extensions in use give it a second outcome instead. Benavoli et al. (2021c) attach a valid or invalid label to each evaluation alongside the preferences, for experiments that sometimes cannot produce an output; C-GLISp learns the probability that a design is feasible and satisfactory from judgments made one design at a time (Zhu et al., 2022); and CrashPBO treats a crash report as a second kind of outcome, reducing crashes by 63% on synthetic benchmarks and validated on three robot platforms (Menn et al., 2026b). The common pattern is two models side by side, one for the utility of designs that work and one for the probability that a design works, combined in the acquisition function much as constrained Bayesian optimization combines an objective with a feasibility model (Section 14.4).
Sources cited in Section 20.4 12
- Benavoli and Azzimonti (2026a) A tutorial on learning from preferences and choices with Gaussian Processes
- Erarslan et al. (2025) Consecutive Preferential Bayesian Optimization
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
- Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
- Cover and Thomas (2006) Elements of Information Theory
- Ou et al. (2023) The Impact of Expertise in the Loop for Exploring Machine Rationality
- Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
- Wu et al. (2025a) Mixed Likelihood Variational Gaussian Processes
- Zhang et al. (2026b) Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
- Menn et al. (2026b) Preferential Bayesian Optimization with Crash Feedback
20.5 Many people #
The crowdsourced line search of Section 20.2.3 took the median of several workers' slider positions, which treats them as noisy measurements of one utility. Koyama and Igarashi (2018) state the assumption plainly: "a common 'general' preference exists that is shared among crowds". They also name its limits: in some domains "crowds from different backgrounds can have clearly different preferences", and "the user's personal preference is not reflected in computation when crowdsourcing is the only source of data". Whenever more than one person answers, the model has to say how their utilities relate.
20.5.1 Three positions and a kernel that spans them #
There are three basic positions. Pool everyone: one utility, and differences between people become noise. Separate everyone: one independent utility per person, which wastes everything the others' answers say. Or share structure: each person's utility is the population's plus a personal deviation. A Gaussian process expresses the third position with one kernel over pairs of (person, design):
The first term is the covariance of a utility shared by everyone; the second is the covariance of an independent deviation for each person , so person 's utility is . If both kernels have the same shape with signal variances and , two people's utilities at the same design have correlation . Setting pools; setting separates. A new person's model then starts from the population's posterior instead of from the prior, and their own answers gradually override it.
Sharing also puts everyone on one scale, the question Section 18.4 left open about comparing utilities across people. Comparisons fix each person's utility only up to a shift and only in units of that person's noise, so a model that adds personal deviations to a common assumes that people are about equally noisy, unless it gives each person a noise scale of their own (inference).
Richer versions share structure without assuming one common utility. The collaborative model of Houlsby et al. (2012) combines a preference kernel with low-dimensional structure shared across users; crowdGPPL writes each person's utility as a weighted combination of a few shared latent functions, a matrix factorization with Gaussian process factors, and scales to thousands of users and items with stochastic variational inference (Simpson and Gurevych, 2020). A 2026 preprint replaces the single utility with a mixture of latent preference archetypes (Dubey et al., 2026).
20.5.2 What population priors buy #
Interactive systems since 2024 have used the population mostly as a prior that gets a new user started. Meta-PO combines PBO with meta-learning, which here means using the Gaussian process models fitted to earlier users to start the search for a new one, and uses the Sequential Gallery interface; in a study of 36 participants in three groups of 12, the iterations needed to reach a satisfactory result fell from 9.54 (standard deviation 2.19) without transfer to 5.86 (standard deviation 1.20) when the earlier users had pursued the same theme, and to 7.41 (standard deviation 1.28) across themes (Li et al., 2025a). The more the new user's goal differs, the less the population helps. HOMI trains an acquisition function represented by a neural network on simulated users before any real user arrives; with 12 participants and a performance objective, it was better than its baselines only at the second and third iterations, and from the sixth iteration on all methods performed at the same level (Liao et al., 2026).
Against this stands consistent evidence that people differ a great deal. When 20 participants with varying design experience judged the same 600 pairs of generated interfaces, they agreed at only 0.25 on a chance-corrected scale where 1 is perfect agreement and 0 is what chance would give (Krippendorff's α; Cohen's κ is the same to two decimals), in a 2026 preprint (Peng et al., 2026). LineCoSpar found that the utilities behind different users' gait preferences differ (Tucker et al., 2020a). A population prior speeds up the first iterations, but no study has shown that it does no harm to a person who differs from the population (inference). Section 32.5 collects the human-factors evidence.
20.5.3 What comparisons can and cannot recover #
Theory adds a warning about pooling. Suppose many people, or one person in varying unobserved situations, answer pairwise comparisons, and a single Bradley-Terry utility is fitted to all of it. Siththaranjan et al. (2024) proved that, with infinite data on a finite set of options, the fitted utility ranks options by their Borda count, the average probability that an option beats a random opponent, and not in general by their expected utility; no method using infinite pairwise data can always recover the expected-utility order. Chapter 21 meets the Borda count again as one of the definitions of a best option. And Chidambaram et al. (2026) showed that if each person makes only one binary comparison, the distribution of preferences in the population cannot be identified at all, while comparisons among three or more options, even incomplete rankings, can identify it under further conditions. For heterogeneous populations, pairs are the weakest form of feedback; Section 29.9 states these results precisely.
Sources cited in Section 20.5 10
- Koyama and Igarashi (2018) Computational Design with Crowds
- Houlsby et al. (2012) Collaborative Gaussian Processes for Preference Learning
- Simpson and Gurevych (2020) Scalable Bayesian preference learning for crowds
- Dubey et al. (2026) Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization
- Li et al. (2025a) Efficient Visual Appearance Optimization by Learning from Prior Preferences
- Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models
- Peng et al. (2026) Efficient Personalization of Generative User Interfaces
- Tucker et al. (2020a) Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits
- Siththaranjan et al. (2024) Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
- Chidambaram et al. (2026) Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
20.6 The interface is part of the model #
Every likelihood in this chapter is a model of a person doing something with an interface. Equation (20.1) models clicking the best of a set; the slider record models dragging and stopping; Equation (20.4) models a third button. Change the interface and the noise changes, the information per answer changes, the effort changes, and sometimes the preference being measured changes too.
The evidence on how people use these interfaces makes the point concrete. With one slider, an iteration of generative image search took 17.2 seconds; with four sliders, 53.4 seconds, though the four-slider version converged in fewer iterations and participants preferred its flexibility; that study had three participants (Chong et al., 2021). In virtual-reality color grading, users preferred to compare two options rather than four (Yuan et al., 2025). A prompt-optimization system that asked only for binary preferences converged faster and with lower workload than its baselines, but was less expressive than baselines that accepted text feedback or manual edits (Li et al., 2026f). Critics of gallery interfaces point out that every option is still chosen by the optimizer and the person can express nothing beyond liking or disliking it (Mo et al., 2024).
The answers also depend on what the person has already seen: in the field deployment described in Section 16.1.1, the experts' ratings were anchored on meshes seen earlier and grew noisier as results improved (Ou et al., 2022). A system that learns from ordinary slider edits instead of explicit queries rests on the assumption that each edit seeks a better design, and its authors note that they have not evaluated when that holds (Koyama and Goto, 2022). Even the parameterization is part of the interface: a 2026 preprint argues that the representation shown to the user is a core part of interaction design rather than a preprocessing step, and with 40 participants found a five-dimensional learned design space better than the original nine procedural parameters (Owaki et al., 2026).
These findings suggest a short list of questions to answer before choosing a query form (inference, from the sections above):
- What does the person do, and what does it cost them? Seconds per answer, and how many answers before fatigue.
- What is recorded? Which likelihood turns the action into evidence, and what information does the record discard?
- What noise does the action add? Perceptual resolution, motor precision, and the effort that pushes people toward an answer that is merely good enough.
- What happens when the person cannot answer? A tie, a skip, a crash, or a forced guess that the model will misread.
- Whose preference is it? One person, a population, or one person in changing situations.
Chapter 32 follows the human side of these questions through the HCI literature, and Section 46.4 turns them into recommendations.
Settled. Choices from sets and rankings have standard likelihoods, Luce's choice model and Plackett-Luce, with Gaussian process versions. Sliders, galleries, and projections beat pairs in their own papers' simulations. In finite-option theory, reporting only the winner of a larger set does not improve on pairs in the worst case, while reporting the top does.
Contested. Whether queries of more than two options help in practice: the two direct experiments disagree, with different acquisition functions. Whether population priors are safe for people unlike the population.
Missing. Controlled human studies comparing query forms under the same acquisition function and budget; acquisition functions derived for slider, plane, and projection queries; an abstain answer distinct from a tie; models of how the interface's own noise (motor, perceptual, effort) enters the likelihood.
Sources cited in Section 20.6 7
- Chong et al. (2021) Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance
- Yuan et al. (2025) Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality
- Li et al. (2026f) Preference-Guided Prompt Optimization for Text-to-Image Generation
- Mo et al. (2024) Cooperative Multi-Objective Bayesian Design Optimization
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Koyama and Goto (2022) BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions
- Owaki et al. (2026) Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs
20.7 Exercises #
(a) Show that the Plackett-Luce model Equation (20.2) with is the Bradley-Terry model. (b) For four options with utilities and , compute the probability that the person ranks them in the order listed, first to fourth, and the probability that they pick the first as the best. (c) Why is the ranking probability so much smaller, and why does that not make the ranking less informative?
Solution
(a) With the product has two factors. The second is a choice from a set of one and equals 1, so
the Bradley-Terry probability.
(b) The exponentials are , , , , with sum 12.11. The first place has probability , which is also the probability of picking the first as best. Given that, the second place has probability , and the third . The ranking has probability .
(c) A ranking is one of possible answers, a winner one of 4, so each particular ranking is less likely. Low probability of each answer is what lets an answer carry more bits: the observed ranking rules out many more alternatives than the observed winner does.
Show that the three probabilities in Equation (20.5) sum to one, and that is largest when . What is its value there for , the setting Bıyık et al. (2019) used in their simulations?
Solution
Write and , so , , and . Then and . Adding gives .
For the maximum, is proportional to , and is smallest at . There and
For this is : when the two options are in fact equal, the modeled person says "about the same" a little under half the time.
Under the kernel Equation (20.7) with and for a common correlation function with , show that the correlation between two different people's utilities at the same design is . A new person arrives and answers nothing. What does the model predict for their utility, and what is its variance compared with that of the population utility ?
Solution
For the covariance is and each variance is , so the correlation is . For a new person , where is independent of all data. Its posterior mean is the posterior mean of , the population's estimate, and its posterior variance is the posterior variance of plus : however much data the population provides, the new person's utility stays uncertain by at least their personal variance until they answer themselves.
In Figure 20.1, choose an end of the slider without moving it. What does the model record? Then argue that the record "chosen beats both ends" discards information that the slider answer contains, and describe one way to keep more of it within Equation (20.1).
Solution
If the person keeps the end , the chosen point coincides with it, the set of three collapses to two, and the record is the pair ; keeping the other end records . In general the person preferred the chosen point to every point on the slider, a continuum of comparisons, while the record keeps two. One way to keep more is to add further points of the segment to the losing set, for example evenly spaced ones, so the choice becomes a choice from options in Equation (20.1). This costs computation, since the latent vector grows with every point, and the added points sit close together on one line, so their comparisons are strongly correlated and add less than their number suggests. Mikkola et al. (2020) take the limit of many points with Gaussian noise instead.
Sources cited in Section 20.7 2
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
Further reading #
- Koyama et al. (2017) introduce sequential line search; the book chapter Koyama and Igarashi (2018) explains the crowdsourcing design, the microtask choices, and the "whose preference" question in plain terms.
- Koyama et al. (2020) extend the slider to a zoomable gallery on a plane and report both simulations and a small user study.
- Mikkola et al. (2020) derive the likelihood of a projective answer as a continuum of comparisons and test it with materials scientists.
- Siivola et al. (2021) and Nguyen et al. (2021) give likelihoods for batch winners, rankings, and ties; Benavoli and Azzimonti (2026a) surveys nine likelihoods for preferences and choices with Gaussian processes.
- Bıyık et al. (2019) add an "About Equal" answer and a stopping rule to active preference learning.
- Houlsby et al. (2012) and Simpson and Gurevych (2020) model many users' preferences with shared structure; Siththaranjan et al. (2024) show what a single Bradley-Terry utility recovers from a mixed population.
- Luce (1959) and Plackett (1975) are the classical sources for choice from sets and for rankings.
References
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §20.1 §20.3
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Cited in §20.3 §20.4 §20.7
- (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika. Cited in §20.1
- (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Cited in §20.3
- (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Cited in §20.5
- (2021). Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance. Computer Graphics Forum. doi:10.1111/cgf.14188. Cited in §20.6
- (2006). Elements of Information Theory. Wiley. Cited in §20.4
- (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. preprint Cited in §20.5
- (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Cited in §20.4
- (2014). The Limits of Attraction. Journal of Marketing Research. Cited in §20.1
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §20.3
- (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. Cited in §20.5
- (1982). Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis. Journal of Consumer Research. Cited in §20.1
- (2022). BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3526113.3545664. Cited in §20.6
- (2018). Computational Design with Crowds. Computational Interaction. Cited in §20.2 §20.5
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Cited in §20.2 §20.3
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Cited in §20.2 §20.3
- (2025a). Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options. NeurIPS 2025. Cited in §20.1
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Cited in §20.3 §20.4
- (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Cited in §20.5
- (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. Cited in §20.6
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Cited in §20.5
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §20.3
- (1959). Individual Choice Behavior: A Theoretical Analysis. Wiley. Cited in §20.1
- (1974). Conditional Logit Analysis of Qualitative Choice Behavior. Frontiers in Econometrics. Cited in §20.1
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Cited in §20.4
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §20.2 §20.3 §20.7
- (2024). Cooperative Multi-Objective Bayesian Design Optimization. ACM Transactions on Interactive Intelligent Systems. doi:10.1145/3657643. Cited in §20.6
- (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Cited in §20.1 §20.3 §20.4
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §20.6
- (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. Cited in §20.4
- (2026). Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs. arXiv. preprint Cited in §20.6
- (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Cited in §20.5
- (1975). The Analysis of Permutations. Journal of the Royal Statistical Society: Series C (Applied Statistics). Cited in §20.1
- (2019b). PAC Battling Bandits in the Plackett-Luce Model. Algorithmic Learning Theory. Cited in §20.1
- (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Cited in §20.1 §20.3
- (2020). Scalable Bayesian preference learning for crowds. Machine Learning. Cited in §20.5
- (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Cited in §20.5
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Cited in §20.3 §20.5
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Cited in §20.3
- (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. preprint Cited in §20.3 §20.4
- (2025). Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality. IEEE Transactions on Visualization and Computer Graphics. Cited in §20.6
- (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Cited in §20.4
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Cited in §20.4