Bayesian Optimization
Part X: Synthesis
中文

Open Problems and the Decisive Experiment

Preferential Bayesian optimization (PBO) finds the design a person likes best by asking them to compare options and fitting a model of their utility to the answers (Chapter 19). The chapters before this one reported what is known about it: the methods, the theory, the studies with people, and what other disciplines say a preference is. The two previous chapters drew two conclusions from that record. A single comparison mixes a stable preference, structured noise, change caused by the questions, and answers with no preference behind them (Chapter 45), and a system built on comparisons should be instrumented and judged against strong simple baselines (Chapter 46).

This chapter lists what is not known. It collects eighteen open problems, and for each one names an experiment and the result that would decide it. The first group concerns preference itself, the second theory and dimension, and the third the choice of methods in practice. One experiment, randomizing the order of queries and retesting a week later, would change our understanding more than any other, and the chapter describes it concretely enough to run. The chapter, and the book, end with a conclusion.

The experiments are proposals. Where we add detail to them (sample sizes, numbers of retest pairs, how an arm is balanced), the detail is our own inference, offered as a starting point for a design rather than as a finding.

47.1 Problems about preference itself #

The first problems decide how a comparison should be understood: as a noisy readout of a fixed utility, or as something the session partly makes. Table 47.1 lists them.

Table 47.1 Open problems about preference itself, each with an experiment and the result that would decide it.
Problem Experiment Result that would decide it
Does the order of queries change the final preference? (Noisy discovery against preference scaffolding) randomize people to query orders and retest a week later; see Section 47.4 see Section 47.4
Does the current best gain value from being chosen again and again? at several delays during the session, insert comparisons of the current best against options rejected earlier; compare likelihoods with and without a choice-history term by their predictive log-likelihood on held-out pairs the model with a choice-history term is significantly better, and the rate at which rejected options win back rises with the delay
How much of the inconsistency within a session is noise, drift, or incompleteness? repeat identical pairs at different lags, insert placebo pairs, and offer a "can't compare" answer agreement that does not depend on the lag means noise dominates, agreement that falls with the lag means drift; placebo pairs give the false-preference rate; "can't compare" answers concentrated on complex pairs support the fourth kind of answer
Does allowing incomplete answers improve the results? within subjects, compare forced choice with an interface that offers "about the same" and "can't compare" with incomplete answers allowed, violations of transitivity fall and the delayed retest of the final design agrees more often; then the extra answers should be the default
Does the hierarchical view of preference apply to design? the same people judge, repeatedly over several sessions, both abstract goals (such as "comfortable") and concrete parameter configurations judgments of abstract goals have significantly higher test-retest reliability than judgments of configurations: the layered model holds; equal reliability: the hierarchical view has no empirical support in design tasks
How does the value of a population prior depend on the domain? in faces or landscapes, and in architecture or artworks, compare a strong population prior, a prior per group, and a weak prior a strong prior wins for the natural domains and a weak or group prior for cultural artifacts: set the strength of the prior by domain

Choice history. Choosing an option raises its value for the chooser (Enisman et al., 2021; Zylberberg et al., 2024), and in a PBO session the current best is the option chosen most often. The test is a model comparison on held-out answers: if a likelihood with a term for recent choices predicts better than one without, and options the person rejected become more likely to win again as time passes, the effect is present in design comparisons, where it has never been measured (Section 45.4).

Noise, drift, and incompleteness. Repeated pairs at different lags separate noise, which does not depend on the lag, from drift, which does. One caution from Figure 46.1: induced change that fades also makes agreement fall with the lag, so a falling curve alone does not identify drift (inference); comparing query orders does (Section 47.4). Placebo pairs, the same design shown twice, estimate how often people express a preference where none can exist (O'Mahony and Wichchukit, 2017). A "can't compare" answer that concentrates on complex pairs would show that the fourth component of Section 45.4 is real in design tasks; in experiments with lotteries and money, 40% to 50% of participants used such an answer when allowed (Nielsen and Rigotti, 2026).

Incomplete answers as a default. Forced choice turns indifference, indecisiveness, and experimentation into the same coin flip (Ok and Tserenjigmid, 2022). If allowing the extra answers reduces transitivity violations (choosing aa over bb and bb over cc but cc over aa) and makes the final design hold up better at a delayed retest, the cost of the extra buttons is repaid (inference).

The hierarchical view. One position in the literature holds that preferences over ultimate goals are stable while preferences over intermediate outcomes are elicited (Section 45.3.2). The evidence for stability comes from self-reported values (Vecchione et al., 2020) and the evidence for construction from choices, so the two levels have never been tested on the same task. Asking the same people, across sessions, to judge both an abstract goal and concrete configurations would test the view where PBO operates.

Population priors by domain. Shared taste is high for faces and landscapes and low for architecture and artworks (Vessel et al., 2018). If a strong population prior helps in the first domains and a weak or group-specific prior in the second, the strength of the prior is a parameter to set per domain (Section 46.3).

Sources cited in Section 47.1 7
  1. Enisman et al. (2021) Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm
  2. Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
  3. O'Mahony and Wichchukit (2017) The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection
  4. Nielsen and Rigotti (2026) Revealed Incomplete Preferences
  5. Ok and Tserenjigmid (2022) Indifference, indecisiveness, experimentation, and stochastic choice
  6. Vecchione et al. (2020) Stability and change of basic personal values in early adolescence: A 2‐year longitudinal study
  7. Vessel et al. (2018) Stronger shared taste for natural aesthetic domains than for artifacts of human culture

47.2 Problems about theory and dimension #

The next three problems decide what the theory guarantees and where the dimension limit of PBO really comes from (Table 47.2). Two quantities from regret theory appear in them (Section 45.1.2): γT\gamma_T, the maximum information gain of the kernel after TT queries, and κ\kappa, an upper bound on the inverse slope of the link function. A third is BB, a bound on the norm of the utility function in the kernel's own function space, which limits how rough the utility may be.

Table 47.2 Open problems about theory and dimension.
Problem Experiment or analysis Result that would decide it
A lower bound for kernelized preference feedback under a logistic or probit link derive minimax lower bounds under Gumbel or normal noise with explicit dependence on κ\kappa; check whether the warm-up constant of MR-LPF hides κ\kappa or exp⁡(B)\exp(B) a lower bound that matches O~(γTT)\tilde O(\sqrt{\gamma_T T}) establishes order optimality; a mismatch means pairwise feedback really is more expensive, or the upper bound can be improved
Order-optimal guarantees for fully sequential algorithms and for the pipeline used in practice analyze the frequentist regret of a Laplace or expectation-propagation posterior with EUBO, or construct a counterexample in which it is inconsistent; try to remove the extra γT\sqrt{\gamma_T} factor of sequential algorithms a sequential algorithm with O~(γTT)\tilde O(\sqrt{\gamma_T T}), a guarantee for the practical pipeline, or a counterexample
Does the dimension ceiling come mainly from the default prior, and does that hold with human comparisons? at budgets of 50, 100, and 200 comparisons, compare qEUBO with the default prior, qEUBO with a dimension-scaled prior, a local method that does not restrict the lengthscale, a linear utility model, and a reduced representation; first in simulation, then with people in 20- to 50-dimensional latent spaces of generative models; record the gradients of the Laplace marginal likelihood with respect to the hyperparameters if the dimension-scaled prior closes most of the gap, the ceiling comes mainly from the defaults; if only the reduced representation works, the limit comes from the budget and the information per query

Lower bounds. All regret results for kernelized preference feedback are upper bounds. The batched algorithm MR-LPF reaches O~(γTT)\tilde O(\sqrt{\gamma_T T}), matching the scalar rate, but needs a warm-up phase whose constant does not depend on TT (Kayal et al., 2025), and the question is whether that constant hides a dependence on κ\kappa or on exp⁡(B)\exp(B), which could be very large. No lower bound exists under a logistic or probit link (Section 29.7).

The pipeline in practice. Theory analyzes elimination, optimism, and Thompson sampling on frequentist kernel estimators; practice runs a Laplace posterior with EUBO (Section 45.1.2). A regret analysis of that pipeline, or a counterexample showing that it can fail to converge, would close the gap between what is proved and what is used.

The dimension ceiling. A common claim is that PBO fails beyond 10 to 20 dimensions. The evidence points instead to a property of the default configuration: PairwiseGP's lengthscale prior does not scale with dimension, so the default kernel treats points in high dimension as nearly unrelated (Section 46.3 gives the arithmetic; see also Section 27.5 and Section 30.3.4). For scalar feedback, dimension-scaled priors repaired this (Hvarfner et al., 2024), but with pairwise feedback and human budgets the repair is untested. The proposed experiment separates two explanations: if the dimension-scaled prior closes most of the gap to a reduced representation, the ceiling was the defaults; if only the reduced representation works, the limit is how little each answer carries. Recording the gradients of the marginal likelihood shows whether the lengthscales are being learned at all. Simulation comes first, so that the experiment with people, in a 20- to 50-dimensional latent space of a generative model, is not confounded by the default prior.

Figure 30.2 runs the scalar version of this experiment on a toy problem: a six-dimensional function hidden among up to 50 inputs, with fixed-scale and dimension-scaled priors and two ways of searching the acquisition function. Its lengthscale panel shows whether the lengthscales are learned, and in those runs the answer is a matter of degree: under the fixed-scale prior the inputs that matter get lengthscales only about three times shorter than the rest, while the dimension-scaled prior switches most of the irrelevant inputs off. The figure's regret curves show why the result must be read together with the acquisition search.

Sources cited in Section 47.2 2
  1. Kayal et al. (2025) Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
  2. Hvarfner et al. (2024) Vanilla Bayesian Optimization Performs Great in High Dimensions

47.3 Problems about practice #

The remaining problems decide which methods to use (Table 47.3). Most of them could be settled by an experiment with people that compares a handful of variants on the same task.

Table 47.3 Open problems about the choice of methods in practice.
Problem Experiment Result that would decide it
Does algorithmic tuning beat manual self-tuning? the same participants and the same time budget, an exoskeleton or a prosthesis, manual self-tuning against pairwise PBO, with the order of conditions balanced and an objective endpoint preregistered metabolic or biomechanical endpoints, delayed consistency of preference, and time; if there is no difference, manual self-tuning should be the default
How large are the differences between acquisition functions with people? the same interface and budget, people randomized to qEUBO, dueling Thompson sampling, the maximally uncertain challenge, and random queries if the differences are smaller than the differences between people and the test-retest noise, the choice of acquisition function is a secondary question in practice
What are response times worth in the closed loop? record response times; in the closed loop, compare a joint likelihood of choices and response times (with the overall value of the pair as a covariate) with the probit likelihood the joint likelihood predicts held-out pairs better, and the stopping times it implies correlate with the agreement of the delayed retest
Can labels from language models replace labels from people? the same task and the same language-in-the-loop pipeline, answered by people and by decision makers simulated with a language model; compare individual consistency, transitivity, position bias, and final designs if individual-level agreement is near the 0.11 to 0.22 found for simulated users in PRISM-X, against 0.57 for the agreement between people's own ratings and rankings, simulated evaluations cannot be extrapolated to people (Kirk et al., 2026) (preprint)
There is no public data set of individual pairwise judgments build a public data set with timestamps, presentation order, response times, repeated pairs, and delayed retests, covering perceptual parameters, aesthetic design, and device tuning once it exists, link functions, noise and drift models, and the gap between simulated and real users can be compared offline
Do EUBO's collapse and the rank-deficient Hessian cost anything with real people? on the same task with people, compare EUBO, EUBO with a connectivity constraint, a diagonal correction scaled by the prior uncertainty, and the exact knowledge gradient if the corrections give significantly better final designs or delayed consistency, the default pipeline needs to change
Does the inference method matter at human noise levels? a closed-loop experiment with people comparing the Laplace approximation, expectation propagation, exact skew Gaussian process sampling, and the hallucination believer no difference in final designs: the debate about inference is secondary in practice; a difference: the default implementation should be replaced
When does active selection beat random selection in collecting preference data for language models? the same model, labels, and compute budget; compare a reward model with explicit epistemic uncertainty and information-directed selection, the implicit reward of direct preference optimization with uncertainty-based selection, and random selection; report win rates and general capability if only the variant with an explicit posterior beats random selection reliably, the gain comes from where the uncertainty is represented
A stopping rule for a pairwise likelihood carry cost-aware stopping over to the pairwise likelihood, and judge it by the agreement of the delayed retest and by the person's satisfaction if the carried-over rule saves comparisons without lowering delayed agreement, it can be the default stopping criterion

Manual against algorithmic tuning. In exoskeleton tuning, people adjusting four hip-timing parameters with a thumbstick reached a 16.6% reduction in metabolic cost in about 11 minutes (Schäfer et al., 2026), of the same order as algorithmic tuning reported in other studies with different devices and controls. No study has compared the two on the same participants against a preregistered endpoint, an outcome fixed before data collection (Section 33.3, Section 24.5). Because people adapt to a device over about 109 minutes of assisted walking (Poggensee and Collins, 2021), the order of the two conditions has to be balanced across participants.

Acquisition functions with people. The candidates in the table are the decision-theoretic qEUBO (Section 19.4), two older heuristics, dueling Thompson sampling, which takes the best option under one random sample of the posterior and pairs it with the option whose duel against it is most uncertain, and the maximally uncertain challenge, which pits the option with the largest posterior mean against the option whose comparison with it is most uncertain in the model's knowledge (Section 28.1), and random queries as a control. We found no study that randomized people to different acquisition functions with the same interface and budget (Section 28.9). If the differences between acquisition functions turn out smaller than the differences between people and the noise of a retest, the work on acquisition functions matters less, in practice, than the observation model (inference).

Response times, simulated users, and data. Response times improved models on real human data (Shvartsman et al., 2024), but they have not been tested in closed-loop optimization. Language-model simulations of users reproduce population-level rankings almost perfectly but agree with individuals at a Kendall τ\tau (a rank correlation from −1-1 to 11) of only 0.11 to 0.22 (Kirk et al., 2026), so whether they can stand in for people in a PBO loop is open (Section 35.2.4). A public data set of individual comparisons, with the timing and order information listed in the table, is the prerequisite for most offline comparisons in this list (Section 31.5).

The default pipeline. EUBO's tendency to collapse toward the estimated maximum, and the rank-deficient Hessian that its isolated pairs cause, are documented in 2026 preprints (Wu and Gardner, 2026; Shao et al., 2026), and the error of the Laplace approximation was measured at very low noise (Takeno et al., 2023). The alternatives in the table are expectation propagation and exact sampling of the skew Gaussian process posterior (Chapter 17), and the hallucination believer, which plugs a single posterior sample in as if it were data (Section 19.3). Whether these failures cost anything with real people, at real noise levels, is the question (Section 27.6, Section 27.4).

Language models and stopping. In online direct preference optimization, a 2026 workshop paper found active selection only negligibly better than random (Oh et al., 2026b); whether an explicit posterior over rewards changes that is open (Section 35.3.5). Stopping rules from scalar Bayesian optimization (Wilson, 2024; Xie et al., 2026) have not been carried over to a pairwise likelihood, and Section 46.6 explains why a person's cost per query is not constant.

Sources cited in Section 47.3 10
  1. Kirk et al. (2026) PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
  2. Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
  3. Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
  4. Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences
  5. Wu and Gardner (2026) Knowledge Gradient for Preference Learning
  6. Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
  7. Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
  8. Oh et al. (2026b) Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
  9. Wilson (2024) Stopping Bayesian Optimization with Probabilistic Regret Bounds
  10. Xie et al. (2026) Cost-aware Stopping for Bayesian Optimization

47.4 The decisive experiment #

Of all the problems above, one would change our understanding most: whether the order of queries changes the final preference. It decides between two readings of every PBO session, noisy discovery, in which repeated comparisons recover a preference that was there before, and preference scaffolding or construction, in which the questions help make the preference they measure (Section 45.3). It also supplies the controls that judging the legitimacy of a system-induced change requires (Section 45.5), and its retests separate noise from drift. As of September 2026, no study has run it.

47.4.1 The design #

In outline, the design is simple. Everyone gets the same candidate space and the same budget of comparisons. Participants are randomly assigned to a sequence of queries chosen by the acquisition function, a random sequence, or a sequence in balanced order. Before the session, each person's attitude toward having their taste changed is recorded. At the end of the session and again a week later, without the system's framing, each person chooses between the final design and options they rejected early. Algorithm 47.1 fills in the details we would choose; every number in it is our suggestion (inference).

Algorithm 47.1 The decisive experiment: randomized query order with a delayed retest (details are our inference)
  1. Task. Choose a perceptual design task with a small, fixed set of rendered candidates, for example 4 to 6 parameters of a visual design or a photo adjustment (Chapter 25), so that final designs from different people can be compared and the task can run online.
  2. Arms. Randomize each participant to one of three arms with the same interface and the same budget, for example 40 comparisons: acquisition (qEUBO or EUBO on PairwiseGP), random (uniformly random pairs), and balanced (every candidate shown equally often, in an order counterbalanced across participants). In every arm, compute the final design and the early favorite with the same model, so that the arms differ only in which questions were asked.
  3. Before the session. Record the participant's attitude toward having their taste changed, and their familiarity with the domain.
  4. During the session. At fixed positions shared by all arms, insert two placebo pairs and a few repeated pairs at different lags. Randomize left and right. Log every answer with its time and response time.
  5. Early favorite. After the 10th answer, record the design with the highest posterior mean.
  6. End-of-session retest. Immediately after the session, in a neutral interface with no labels and no history, show the final design against each of about 8 designs the participant rejected in the first 10 answers, in random order and position, mixed with filler pairs. Ask whether they endorse the final design.
  7. Delayed retest. One week later, repeat step 6 with the same pairs.
  8. Analysis. Preregister the hypotheses, outcomes, and analysis below before collecting data.

47.4.2 What decides it #

The decision criterion has two halves. If the final designs of the acquisition arm lie closer to their own early posterior mean, and their delayed retest agrees less often than in the random arm, then change induced by the queries is a substantive component of the answers. If the distribution of final designs and the delayed agreement do not differ between the arms, noisy discovery is enough to explain the data.

Our analysis of the model in Figure 45.2 suggests a refinement (inference). In that model, the second half of the criterion, comparing the delayed agreement of the two arms directly, can fail even when induced change is present and fades. At the figure's default settings, the acquisition arm's final designs are still preferred more often a week later than the random arm's (78% against 73%), because the acquisition arm reaches a higher agreement at the end of the session (88% against 76%) and part of that lead survives. The signature that holds up is the drop in agreement from the end of the session to the delayed retest, compared across arms: 10 percentage points in the acquisition arm against 3 in the random arm, while with induced change set to zero both drops vanish. A reader can reproduce this by setting the figure's sliders and reading its retest panel.

The model also shows the limit of the drop. When the induced change lasts (the Lasting share slider at 1), neither arm drops, and only the first half of the criterion, final designs close to the early favorite and a different distribution of final designs across arms, still detects it (Exercise 45.1). We would therefore preregister three contrasts (inference):

  • Primary: the mean drop from the end-of-session retest to the delayed retest, acquisition arm against random arm (a difference in differences).
  • Secondary: the distance between the final design and the early favorite, and the distribution of final designs, compared across arms; and the direct comparison of delayed agreement between arms.
  • Diagnostic: agreement on repeated pairs by lag and the false-preference rate on placebo pairs, pooled over arms, which separate noise and drift as in Figure 46.1.
Table 47.4 Outcomes of the decisive experiment and how to read them (inference).
Pattern Reading
No drop in either arm, same distribution of final designs noisy discovery explains the data
A larger drop in the acquisition arm than in the random arm change induced by the queries, at least partly transient
No drop, but the acquisition arm's final designs sit closer to the early favorite and differ in distribution lasting change induced by the queries; whether it is acceptable is a question of legitimacy, not measurement
Equal drops in all arms drift or fatigue that does not depend on the queries
Low agreement at both retests in all arms, many "can't compare" answers on complex pairs noise and incompleteness dominate

47.4.3 How many people #

The experiment is cheap in equipment and expensive in participants, because the primary outcome is a small difference between two differences of proportions. Figure 47.1 gives a planning approximation: a person's measured drop varies because of the binomial noise of a few retest choices at each time and because people differ, and a standard two-sample test then needs the number of participants per arm that the figure marks.

0.00.20.40.60.81.0power103010030010003000participants per arm80% power161 per arm · 483 in three armsspread of one person's measured drop: SD 0.22 = retest noise 0.20 and differences between people 0.10
0.00.20.40.60.81.0power103010030010003000participants per arm80% power161 per armSD of one person's drop: 0.22
Figure 47.1 A planning approximation for the decisive experiment, not a result. The curve is the power of a two-sided test at the 5% level to detect the chosen difference in the drop of retest agreement between the acquisition and random arms, against the number of participants per arm. A person's measured drop has variance equal to the spread of true drops between people plus the binomial noise of their retest choices at the two times; the approximation ignores clustering and any correlation between a person's two retests. The default difference of 7 percentage points is the one in the simulation of Figure 45.2 at its default settings; the other defaults are illustrative.

At the defaults, a difference of 7 percentage points in the drop, 8 retest pairs per person, agreement around 80%, and a spread of 0.1 between people, about 161 participants per arm reach 80% power, 483 in three arms. Doubling the retest pairs to 16 cuts that to 97 per arm; a difference of 10 points with 16 pairs needs 48. The number of retest pairs is the cheapest lever, since each pair costs seconds while each participant costs a session and a return visit a week later (inference). Numbers of this size are within reach of an online study with a visual task, and out of reach of a laboratory study with a device, which is one reason to run the experiment first where it is cheap.

47.4.4 Practical notes #

Keep the task low-stakes. The experiment deliberately includes an arm that may shape preferences more than others. A design task without consequences beyond the session, informed consent that explains that the system may influence choices, and a debriefing after the delayed retest keep it within ordinary research ethics (inference).

Extensions. The same design answers other problems in this chapter with one more arm each: an arm with manual self-tuning answers whether the algorithm is needed at all, arms with different acquisition functions answer how much they differ with people, and an arm with the extra answers "about the same" and "can't compare" answers whether incomplete answers should be the default.

Why it has not been run. It needs no new method, only a commitment to measuring the system's effect on the people it optimizes for. The two retests, the random arm, and the attitude question are the same controls that Algorithm 46.2 recommends for every session, so a study that already follows those recommendations has most of the decisive experiment built in.

47.5 Settled, contested, missing #

Research status Settled, contested, missing

Settled. None of the eighteen problems. What is settled is the ground they stand on: choosing changes preferences by a considerable amount; a forced choice hides incomplete preferences; regret bounds for preference feedback are upper bounds only; the default PairwiseGP prior does not scale with dimension; and a disconnected comparison graph leaves relative utilities undetermined.

Contested. Whether EUBO's collapse and the rank-deficient Hessian matter with people (2026 preprints); whether inference errors matter at human noise levels (two groups disagree about realistic tests); whether random selection is hard to beat beyond the online preference optimization of language models (a workshop paper).

Missing. Every experiment with people in this chapter: randomized query order with a delayed retest; manual against algorithmic tuning on the same participants with a preregistered endpoint; acquisition functions compared on people; dimension-scaled priors tested with human comparisons; and a public data set of individual pairwise judgments. On the theoretical side, a lower bound under a logistic or probit link.

47.6 Outlook: where the problem is growing #

The open problems above are about what is not yet known. This section is about where the research is worth doing, and it is the book's own view rather than a summary of evidence: every paragraph in it could carry the mark (inference), so it is marked once here.

It helps to separate two questions. Bayesian optimization and its preferential variant are answers to one problem: finding what is good when each evaluation is expensive, and, in the preferential case, when the only instrument is a person's judgment. The first question is whether the algorithm family has much further to go. The second is whether the problem it answers is becoming more or less important.

The algorithm family. Its future is modest. The core is mature: decision-theoretic acquisition functions with a multi-option form (Section 19.4), a default pipeline in widely used software (Section 31.1.1), and regret theory that has caught up with scalar feedback at the level of upper bounds (Section 29.4). The evidence with people is small: the median user study in the evidence map of Figure 32.3 has 12 participants, and simpler methods often do as well (Section 45.1.5). Another acquisition function is unlikely to change that picture. Better measurement could.

The problem. It is growing. Systems increasingly generate, adapt, and personalize, and each of those needs to learn what one particular person wants from a handful of judgments. The directions below are where that problem meets an opening, ordered roughly from the measurement questions at the center of this book outward. Each names what exists, what is missing, and the combination of skills it rewards.

47.6.1 Measurement science for preferences #

The niche closest to this book's argument. The observation model, probit or logistic noise of constant size, has never been compared with an alternative on human data; noise that depends on how hard a comparison is, response times, "can't tell" answers, and drift over a session are all open (Section 47.1, Section 27.2). The decisive experiment of Section 47.4, randomized query order with a retest a week later, would separate a preference that was found from one that was shaped, and nobody has run it. And no public data set of individual pairwise judgments exists (Section 31.5), so almost every offline comparison of methods runs on simulated people. The work is cheap, mostly participant time with existing interfaces, and has high leverage, because every method in this book sits on top of the observation model. It rewards someone who can design a careful experiment with people and also write the model: human-computer interaction or psychophysics together with probabilistic machine learning.

47.6.2 Personalizing generative models #

A foundation model is a strong prior over what people in general like and a weak one about what a specific person likes. A handful of well-chosen comparisons on top of such a prior is exactly the preferential setup, with a far better prior than a Gaussian process on raw parameters. The first systems exist: GimmBO searches the merging weights of 20 to 30 diffusion adapters (Liu et al., 2026b), MultiBO shows several images per round (Rajagopalan et al., 2026), and APPO hides the text prompt and asks only for binary preferences between images while a language model rewrites the prompt (Li et al., 2026f) (Section 32.1.6). Missing are a surrogate that uses the generative model itself rather than a Gaussian process over its latent space (Section 30.5), tests with people in the 20 to 50 dimensions such spaces have (Section 30.3.4), and an evaluation that separates what the person specifically likes from what the population likes (Section 32.5). It rewards generative modeling together with Bayesian optimization and user studies.

47.6.3 Choosing which comparisons to collect for language models #

Reward models and direct preference optimization learn from pairwise comparisons at scale, and they face the questions this book studies: which pair to ask about next, and when to stop (Section 35.3.3, Section 35.3.4). The setting differs in ways that matter: millions of comparisons instead of dozens, many raters instead of one person, and a population preference as the target (Section 35.6.2). Whether choosing pairs actively beats random pairs at that scale is itself unsettled (Section 35.3.5), and disagreement between raters may be signal rather than noise, the problem social choice studies (Section 40.7). This is where most of the funding and attention are. It rewards large-scale machine-learning engineering together with the theory of active learning, and it gains from the measurement work above, since raters are people too.

47.6.4 Devices on the body #

Exoskeletons, prostheses, hearing aids, and neurostimulation share the conditions under which preferential optimization makes most sense: comparisons are natural, objective measures are slow or expensive, and people differ more than any population default allows. The strongest real-world evidence in this book comes from here. Bayesian optimization tuned two timing parameters of a soft exosuit in 21.4 ± 1.0 minutes (Ding et al., 2018); wearers tuning four parameters themselves with a thumbstick remote control took 10.9 minutes (Schäfer et al., 2026); preset selection for over-the-counter hearing aids has been posed as a dueling bandit (Vyas et al., 2022); a variant of safe optimization that takes preference feedback was applied to spinal cord stimulation (Sui et al., 2018b); and a thesis tuned the encoders of retinal prostheses with people in the loop (Fauvel, 2021) (Chapter 33). Missing are a head-to-head comparison of manual and algorithmic tuning on the same participants (Section 33.3), exploration that stays safe for the body, and studies that follow a person for weeks while their body and preference adapt. It rewards biomechanics or clinical research together with safe and non-stationary Bayesian optimization.

47.6.5 Simulated people and model judges #

Because studies with people are expensive, nearly every comparison of methods uses simulated people (Section 31.7), and language models now act as judges and stand-in users. One measurement shows the risk: simulated users reproduced population-level rankings almost perfectly but agreed with individuals at a Kendall τ\tau of only 0.11 to 0.22 (Kirk et al., 2026) (Section 35.2.4). A simulator validated against real comparison data, with order effects and drift, would make offline benchmarks informative; an unvalidated one makes them circular. It rewards user modeling together with evaluation methodology, and it depends on the public data set the first direction would produce.

47.6.6 Science and engineering with experts in the loop #

Self-driving laboratories now report how much faster an optimizer reaches a target than a reference strategy (Adesiji et al., 2026). Expert judgment helped where the goal had no sensor, as in printing objects with subjective qualities (Deneault et al., 2025), and lost to the optimizer where yield could be measured (Shields et al., 2021) (Section 34.2). Missing are models that combine measured objectives with an expert's pairwise preferences, queries that treat the expert's time as a cost, and studies with more than one expert in the loop. It rewards domain science together with multi-objective and cost-aware Bayesian optimization.

47.6.7 Many people, one setting #

A thermostat in a shared office, a recommender, or a design meant for a population is tuned by the preferences of many people at once (Section 34.1.1). Aggregation, fairness, and strategic answers are questions social choice has studied for decades (Section 40.7), and a prior learned from earlier people in the same domain could cut the number of questions for each new one (Section 47.1). It rewards mechanism design or social choice together with preference learning.

47.6.8 Theory for the pipeline people deploy #

All regret results for kernelized preference feedback are upper bounds; there is no lower bound under a logistic or probit link, and the algorithms that theory analyzes, elimination and optimism, are not the decision-theoretic acquisition with a Laplace approximation that software runs (Section 47.2). Preferences that drift during a session have almost no theory. It rewards learning theory, with enough contact with practice to analyze what is actually deployed.

47.6.9 Signals beyond the click #

A comparison records more than which option won: how long the person took, and how much they hesitated. A drift-diffusion model links choice and time (Section 39.3), and response times improved models on real human data (Shvartsman et al., 2024), but they have not been used to choose the next query. Eye movements and physiological signals are further candidates. The open question is which of these signals are robust across people and interfaces, and whether they improve the choice of queries and not only the fit. It rewards cognitive modeling together with Bayesian optimization.

47.6.10 The intervention as a design problem #

If a session can change a preference, a system should be able to detect when it does, and a study should state the conditions under which that change is acceptable (Section 45.5.2, Section 41.2). Measuring preference change within a session, stopping rules that account for it, and interfaces that show a person how their answers moved are technical problems, not only ethical ones. It rewards human-computer interaction together with statistics and moral philosophy.

Table 47.5 Research directions where the problem of learning what is good from few judgments is growing: what is missing, and what each rewards.
Direction Most important missing piece Rewards Start with
Measurement science observation models tested on human data; the randomized retest experiments with people + probabilistic modeling Section 47.4
Generative models the generative model as the surrogate; tests in 20 to 50 dimensions generative modeling + BO + user studies Section 32.1.6
Language-model comparisons whether active pair selection helps at scale large-scale ML + active learning Section 35.3.3
Devices on the body manual against algorithmic tuning, same people biomechanics or clinical research + safe BO Section 33.3
Simulated people simulators validated on individual data user modeling + evaluation Section 31.7
Experts in the loop measured objectives combined with expert preferences domain science + multi-objective BO Section 34.2
Many people aggregation and fairness for shared settings social choice + preference learning Section 40.7
Theory a lower bound under a logistic or probit link learning theory Section 47.2
Signals beyond the click response times used to choose queries cognitive modeling + BO Section 39.3
The intervention measuring preference change within a session HCI + statistics + ethics Section 45.5.2

What the directions share is the book's argument. The bottleneck has moved from algorithms to measurement, and most of the valuable work sits where a careful experiment with people meets a well-specified model. People who can do both are scarce, which is the opportunity.

Sources cited in Section 47.6 13
  1. Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
  2. Rajagopalan et al. (2026) Personalized Image Generation via Human-in-the-loop Bayesian Optimization
  3. Li et al. (2026f) Preference-Guided Prompt Optimization for Text-to-Image Generation
  4. Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
  5. Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
  6. Vyas et al. (2022) Personalizing over-the-counter hearing aids using pairwise comparisons
  7. Sui et al. (2018b) Stagewise Safe Bayesian Optimization with Gaussian Processes
  8. Fauvel (2021) Human-in-the-loop optimization of retinal prostheses encoders
  9. Kirk et al. (2026) PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
  10. Adesiji et al. (2026) Benchmarking self-driving labs
  11. Deneault et al. (2025) Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities
  12. Shields et al. (2021) Bayesian reaction optimization as a tool for chemical synthesis
  13. Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences

47.7 Conclusion #

Between 2017 and 2026 the largest change in PBO was not a new methodological paradigm but a shift in where the attention went. Acquisition functions gained a decision-theoretic foundation, regret theory caught up with scalar feedback at the level of upper bounds, and software fixed a default pipeline. At the same time, every layer of that default was found to rest on an untested assumption: the lengthscale prior does not scale with dimension, the Laplace approximation is ill-conditioned when the comparison graph is disconnected, and the probit link with constant noise has never been compared with an alternative on human data. The evidence gives an order of work: first get the observation model and the representation shown to the person right, then choose the acquisition function.

About preference itself, the evidence since 2017 supports neither pure discovery nor pure construction. A single pairwise judgment contains a stable component, structured evaluation noise, change caused by the query, and incomplete or deliberately randomized answers, and which of them dominates depends on the domain, the similarity of the options, and the length of the session. The practical consequence is that the output of a PBO session is both an estimate and an intervention. A study that does not randomize the order of queries or retest after a delay cannot tell whether what converged was the preference or the system's effect on the person, and because the controls needed to tell them apart are the controls needed to judge whether that effect was legitimate, the methodological question and the normative one can be answered in the same experiment (inference).

Three positions in the literature come out of this with qualified support. The hierarchical view of preference is supported where it says that stability lies at the broad, abstract, self-reported level, and not supported where it predicts that iterated querying converges on an underlying preference (Section 45.3.2). Preference scaffolding is partly supported as a description of what PBO does, but as a normative position it needs conditions of legitimacy, and endorsement after the fact cannot supply them on its own (Section 45.3.3). And the common claim that PBO stops working beyond 10 to 20 dimensions describes the default configuration, not a limit of preference feedback (Section 30.3.4).

Four pieces of work are the most likely to change these conclusions: randomized query order with a delayed retest; manual self-tuning against algorithmic tuning on the same participants; dimension-scaled priors tested with human comparisons on 20- to 50-dimensional representations; and a public data set of individual pairwise judgments. The first two can be done with existing devices and interfaces, and their main cost is participants' time; the third needs simulation first, to rule out the default prior as a confound; the fourth is the precondition for most of the offline comparisons in this chapter. On the theoretical side, the most important gap is a lower bound under a logistic or probit link.

Until these results exist, the evidence supports a modest practice: treat PBO as one of several baselines, instrument every session so that it measures the person as well as the design, and, in every study, measure what the system does to the people it learns from.

47.8 Exercises #

Exercise 47.1

Using Figure 47.1, a team can afford 300 participants in total across three arms and expects a difference of 7 percentage points in the drop. How many retest pairs per person do they need for 80% power, assuming agreement around 80% and a spread of 0.1 between people? What if the spread between people is 0.15?

Solution

With 100 participants per arm, the variance s2s^2 of a person's drop must satisfy 2(1.96+0.84)2s2/0.072≤1002 (1.96 + 0.84)^2 s^2 / 0.07^2 \le 100, so s2≤0.0312s^2 \le 0.0312. With a spread of 0.1 between people, the binomial part, 2⋅0.8⋅0.2/k=0.32/k2 \cdot 0.8 \cdot 0.2 / k = 0.32 / k, must be at most 0.0312−0.01=0.02120.0312 - 0.01 = 0.0212, so k≥15.1k \ge 15.1: 16 retest pairs per person (the figure shows 97 per arm at 16 pairs). With a spread of 0.15, the binomial part must be at most 0.0312−0.0225=0.00870.0312 - 0.0225 = 0.0087, so k≥0.32/0.0087≈37k \ge 0.32 / 0.0087 \approx 37, beyond the figure's range. Once differences between people dominate, more retest pairs barely help, and the team needs more participants or a larger expected effect.

Exercise 47.2

Sketch the experiment for "Does algorithmic tuning beat manual self-tuning?" on an ankle exoskeleton with two parameters, peak torque and its timing. What are the arms, the order, the primary endpoint, and the main threat to validity?

Solution

A within-subject design: each participant does both conditions, manual self-tuning (adjusting the two parameters directly, as in self-tuning studies of ankle torque) and pairwise PBO, with the same time budget, in an order counterbalanced across participants. The primary endpoint, preregistered, could be the metabolic cost of walking with each final setting, measured in a blinded validation block at the end of each condition; secondary endpoints are the time to converge and the delayed consistency of each person's preference. The main threat is adaptation: people take on the order of 100 minutes of assisted walking to become expert users (Poggensee and Collins, 2021), so a familiarization phase before either condition, and counterbalancing, are needed to keep learning from favoring whichever condition comes second.

Sources cited in Section 47.8 1
  1. Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance

Further reading #

References

  1. Adesiji, A. D., Wang, J., Kuo, C.-S., and Brown, K. A. (2026). Benchmarking self-driving labs. Digital Discovery. Cited in §47.6
  2. Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning.
  3. Dean, S., and Morgenstern, J. (2022). Preference Dynamics Under Personalized Recommendations. EC 2022.
  4. Deneault, J. R., Kim, W., Kim, J., Gu, Y., Chang, J., Maruyama, B., Myung, J. I., and Pitt, M. A. (2025). Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities. Digital Discovery. Cited in §47.6
  5. Ding, Y., Kim, M., Kuindersma, S., and Walsh, C. J. (2018). Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking. Science Robotics. Cited in §47.6
  6. Enisman, M., Shpitzer, H., and Kleiman, T. (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Cited in §47.1
  7. Fauvel, T. (2021). Human-in-the-loop optimization of retinal prostheses encoders. Sorbonne Université. thesis Cited in §47.6
  8. Hvarfner, C., Hellsten, E. O., and Nardi, L. (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. Cited in §47.2
  9. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Cited in §47.2
  10. Keswani, V., Cousins, C., Nguyen, B., Conitzer, V., Heidari, H., Borg, J. S., and Sinnott-Armstrong, W. (2026). Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback. AAAI.
  11. Kirk, H. R., Leqi, L., Zeng, F., Davidson, H., Vidgen, B., Summerfield, C., and Hale, S. A. (2026). PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users. arXiv. preprint Cited in §47.3 §47.6
  12. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics.
  13. Li, Z., Liao, Y.-C., and Holz, C. (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. Cited in §47.6
  14. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §47.6
  15. Nielsen, K., and Rigotti, L. (2026). Revealed Incomplete Preferences. Working paper (author's website). working paper Cited in §47.1
  16. O'Mahony, M., and Wichchukit, S. (2017). The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection. Trends in Food Science & Technology. Cited in §47.1
  17. Oh, G., Lee, J., Park, J., Yu, Y., Bae, W., and Noh, J. (2026b). Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs. ICLR 2026 Workshop: I Can't Believe It's Not Better (ICBINB). workshop paper Cited in §47.3
  18. Ok, E. A., and Tserenjigmid, G. (2022). Indifference, indecisiveness, experimentation, and stochastic choice. Theoretical Economics. Cited in §47.1
  19. Poggensee, K. L., and Collins, S. H. (2021). How adaptation, training, and customization contribute to benefits from exoskeleton assistance. Science Robotics. Cited in §47.3 §47.8
  20. Rajagopalan, R., Dutta, D., Wei, Y.-L., and Roy Choudhury, R. (2026). Personalized Image Generation via Human-in-the-loop Bayesian Optimization. International Conference on Machine Learning. Cited in §47.6
  21. Schäfer, N., Zhao, G., Li, B., Kupnik, M., Seyfarth, A., Beckerle, P., and Grimmer, M. (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §47.3 §47.6
  22. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §47.3
  23. Shields, B. J., Stevens, J., Li, J., Parasram, M., Damani, F., Alvarado, J. I. M., … Doyle, A. G. (2021). Bayesian reaction optimization as a tool for chemical synthesis. Nature. Cited in §47.6
  24. Shvartsman, M., Letham, B., Bakshy, E., and Keeley, S. (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Cited in §47.3 §47.6
  25. Sui, Y., Zhuang, V., Burdick, J., and Yue, Y. (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. Cited in §47.6
  26. Takeno, S., Nomura, M., and Karasuyama, M. (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §47.3
  27. Vecchione, M., Schwartz, S. H., Davidov, E., Cieciuch, J., Alessandri, G., and Marsicano, G. (2020). Stability and change of basic personal values in early adolescence: A 2‐year longitudinal study. Journal of Personality. Cited in §47.1
  28. Vessel, E. A., Maurer, N., Denker, A. H., and Starr, G. G. (2018). Stronger shared taste for natural aesthetic domains than for artifacts of human culture. Cognition. Cited in §47.1
  29. Vyas, D., Brummet, R., Anwar, Y., Jensen, J., Jorgensen, E., Wu, Y.-H., and Chipara, O. (2022). Personalizing over-the-counter hearing aids using pairwise comparisons. Smart Health. doi:10.1016/j.smhl.2021.100231. Cited in §47.6
  30. Wilson, J. T. (2024). Stopping Bayesian Optimization with Probabilistic Regret Bounds. NeurIPS 2024. Cited in §47.3
  31. Wu, K., and Gardner, J. R. (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Cited in §47.3
  32. Xie, Q., Cai, L., Terenin, A., Frazier, P. I., and Scully, Z. (2026). Cost-aware Stopping for Bayesian Optimization. International Conference on Machine Learning. Cited in §47.3
  33. Zylberberg, A., Bakkour, A., Shohamy, D., and Shadlen, M. N. (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Cited in §47.1