Building and Evaluating a Preferential Optimization System
Preferential Bayesian optimization (PBO) searches for the design a person likes best by asking them to compare options and fitting a model of their utility to the answers (Chapter 19). The previous chapter concluded that a single answer mixes a stable preference with structured noise, change caused by the questions, and answers with no preference behind them, so that a session is both an estimate of a preference and an intervention on it, and that the largest errors come from what the model assumes about the person (Chapter 45).
This chapter turns that analysis into advice for someone who has to build or evaluate a preferential optimization system: three questions that decide whether PBO is the right tool, the parts of the system one at a time, and the situations in which a simpler method should be the default. Every recommendation comes from evidence reported earlier in the book or from our inference from it, and none has been tested as a whole in PBO. Each table also says whether a recommendation depends on the Gaussian process and PBO framework. Those that do not apply to any preference model, including the simple baselines of Section 46.9; those that do are repairs inside the framework, and say nothing about whether the framework was the right choice.
46.1 Three questions before choosing PBO #
Faced with a new problem of personalization or design with a person in the loop, we suggest answering three questions before deciding to use PBO (inference).
- Is there a numerical objective that can be measured? If there is, optimize it with ordinary Bayesian optimization (Chapter 11) or another method suited to it, and use preferences only as a constraint or an auxiliary signal. An expert's pairwise input pays off only when their judgment carries information the surrogate lacks (Section 36.7). Tuning a classifier's validation error (Chapter 22) and maximizing a reaction's yield (Chapter 23) are problems of this kind.
- Can the dimension exposed to the person be brought to about 10 or fewer? If not, design the representation first. With budgets of tens to two hundred comparisons, the limit comes from how little each answer can carry, at most one bit, and from the budget, not from the surrogate (inference).
- With few parameters and a flat optimum, would manual self-tuning or a coarse search already do? If it might, it is the first baseline, and PBO has to beat it (Section 24.5).
When all three questions point toward PBO, qEUBO with BoTorch's PairwiseGP is
a reasonable default (Section 19.4, Section 18.6), a choice that comes from the
software ecosystem rather than from a comparison with simpler methods. The
session should be instrumented as in Algorithm 46.2: the errors caused by the
observation model and by the person's behavior exceed the differences between
acquisition functions (Section 45.1), and these additions both estimate
those errors and supply the controls needed to tell discovery of a preference
from shaping it (Section 45.5).
- Before the session, record the person's attitude toward having their taste changed by the system.
- Decide in advance a fixed share of queries that the acquisition function does not choose: random pairs, or pairs in a balanced order. Randomize which option appears left or right, or where it sits in a gallery.
- Offer two answers besides "A" and "B": "about the same" and "can't compare".
- Insert a few placebo pairs, the same design shown twice, and a few repeated pairs at different lags.
- Record the response time of every answer.
- Label the system's proposals as such, show the comparison history accurately, and let the person reset or reject the current best.
- At the end, retest the final design against designs the person rejected earlier, without the system's framing, and ask whether they endorse it.
- At the next session, a week or so later, repeat that retest.
Steps 3 to 5 are justified in Section 46.2, step 2 in Section 46.4, step 4 again in Section 46.5, steps 7 and 8 in Section 46.6, and steps 1, 6, 7, and 8 in Section 46.8.
A note on dimension. A comparison carries at most one bit (Section 16.6), so 100 comparisons over 20 parameters carry at most five bits per parameter, and usually far fewer, because answers are noisy and many queries are redundant (inference). Where PBO has worked in more than a handful of dimensions, the dimension the person faced had been reduced by design, to a 10-dimensional body-shape space of a generative human model or the merging weights of 20 to 30 diffusion adapters (Koyama et al., 2020; Liu et al., 2026b), or to a 5-dimensional feasibility-aware latent space (Table 46.8).
Sources cited in Section 46.1 2
- Koyama et al. (2020) Sequential Gallery for Interactive Visual Design Optimization
- Liu et al. (2026b) GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
46.2 Observation model #
The observation model carries the largest error (Section 45.1.3). Table 46.1 lists five recommendations for it; none depends on the Gaussian process framework.
| Recommendation | Basis | Depends on the GP and PBO framework? |
|---|---|---|
| Offer "about the same" and "can't compare" besides the two choices; model them as an indifference threshold and as incompleteness (or a mixture parameter), and never read them as a 50:50 answer | Nielsen and Rigotti (2026) (working paper); Ok and Tserenjigmid (2022); Erarslan et al. (2025) (preprint) | no |
| Insert a few placebo pairs (the same design twice) in every session to estimate the false-preference rate and calibrate the noise or lapse term | O'Mahony and Wichchukit (2017) | no |
| Let the noise scale vary with difficulty, similarity, and type of feedback (pairwise answers, sliders, ratings); do not fix a low default or one scale for everyone | Shen et al. (2025b); Ghosal et al. (2023); Keffert and Schweizer (2024) (preprint) | no |
| Record response times; in a joint likelihood, include the pair's overall value as a covariate | Shvartsman et al. (2024); Sawarni et al. (2025); Shevlin et al. (2022) | recording no; the joint likelihood has a Gaussian process version |
| Add a choice-history term that raises the value of recently chosen options; estimate behavioral parameters such as loss aversion rather than fixing them | Zylberberg et al. (2024); Brown et al. (2024) | no |
Two extra answers. A forced choice turns indifference, indecisiveness, and experimentation (Ok and Tserenjigmid, 2022) into a coin flip that the model reads as a weak preference, and when allowed, many people say they cannot compare (Table 45.2). An "about the same" answer has a standard model, a just-noticeable-difference threshold, and a 2025 preprint reports that a method which models it clearly outperforms standard preferential baselines once 10% to 20% of comparisons are indifferent, on synthetic benchmarks (Erarslan et al., 2025) (Section 20.4).
Placebo pairs. Sensory science found that consumers report preferences between identical samples, and made pairs of identical samples a standard control condition (O'Mahony and Wichchukit, 2017). In PBO, the rate at which a person prefers one of two identical renders, identical down to the pixel, estimates how many answers carry no preference at all and can calibrate a lapse term (inference).
Structured noise. In estimates of risk aversion from a representative
sample, the standard expected-utility random utility model gave distorted
estimates when everyone shared one noise scale and not when each person had
their own, in a 2024 preprint (Keffert and Schweizer, 2024). PairwiseGP fixes the
noise at 1 and lets the kernel's amplitude (BoTorch's output scale) stand in
for it (Section 27.5).
Response times are free information about how far apart the options felt. A joint likelihood of choices and drift-diffusion response times exists for Gaussian process models (Shvartsman et al., 2024), and a loss built on the EZ-diffusion model reduces the error of preference learning with linear rewards from exponential to polynomial growth in the magnitude of the reward (Sawarni et al., 2025). But decisions between two high-value options are mostly faster and more accurate, not slower (Shevlin et al., 2022), so late in a session a quick answer does not mean a large difference unless the pair's overall value is in the model (inference).
Choice history. The current best, the most-compared option, gains most from choice-induced revaluation (Section 45.4), and the loss-aversion coefficient has a meta-analytic mean of 1.955 but is about 1.07 when gains and losses are symmetric and unsorted (Table 45.1), so both belong in the likelihood as estimated terms, not fixed values.
Sources cited in Section 46.2 12
- Nielsen and Rigotti (2026) Revealed Incomplete Preferences
- Ok and Tserenjigmid (2022) Indifference, indecisiveness, experimentation, and stochastic choice
- Erarslan et al. (2025) Consecutive Preferential Bayesian Optimization
- O'Mahony and Wichchukit (2017) The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection
- Shen et al. (2025b) Early versus late noise differentially enhances or degrades context-dependent choice
- Ghosal et al. (2023) The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
- Keffert and Schweizer (2024) Stochastic Monotonicity and Random Utility Models: The Good and The Ugly
- Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences
- Sawarni et al. (2025) Preference Learning with Response Time: Robust Losses and Guarantees
- Shevlin et al. (2022) High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity
- Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
- Brown et al. (2024) Meta-analysis of Empirical Estimates of Loss Aversion
46.3 Surrogate #
The surrogate is the prior over the utility function. Table 46.2 lists three recommendations, of which only the last is purely a matter of the Gaussian process framework.
| Recommendation | Basis | Depends on the GP and PBO framework? |
|---|---|---|
| Set the weight of a population prior by domain; expect a population prior to help only in the first iterations, and monitor how far the individual departs from it | Vessel et al. (2018); a prior from user models was significantly better only at iterations 2 and 3 (Liao et al., 2026) | no |
In high dimension, first reduce the dimension the person sees, then consider a linear utility model; if you keep PairwiseGP, switch to a dimension-scaled lengthscale prior and validate it yourself |
Owaki et al. (2026) (preprint); Doumont et al. (2026); Hvarfner et al. (2024) | the first two steps no; the third is a repair inside the framework |
| At low noise, do not use the Laplace approximation; use expectation propagation or exact skew Gaussian process sampling | Takeno et al. (2023); Benavoli et al. (2021c) | yes |
Population priors. Shared taste varies by domain (Table 45.3), and in interactive design a prior pretrained on models of earlier users was significantly better only at the second and third iterations; from the sixth on, all methods were comparable (Liao et al., 2026). A population prior buys a better start, not a better result, and its weight should fall as the person's own answers accumulate (inference; Section 20.5).
High dimension. The default PairwiseGP lengthscale prior is far shorter
than the typical distance between designs in high dimension (about 0.52
against 2.9 at ), so the default model treats every comparison as
nearly uninformative about every other point (Section 27.5,
Section 30.3.2). For scalar feedback, priors that scale with dimension
repaired this (Hvarfner et al., 2024), and a spherical input mapping with
Bayesian linear regression reached the state of the art (Doumont et al., 2026).
Neither has been tested with pairwise feedback from people, so reduce the
dimension the person sees first, which helps under any model. Figure 30.2
adds a second reason: on a toy problem with scalar feedback, how the
acquisition function is searched decided more of the outcome within 80
evaluations than the lengthscale prior did (Section 30.1.2).
Inference. At very low noise, the Laplace approximation misplaces the posterior and expectation propagation distorts duel probabilities near 0 or 1 (Takeno et al., 2023); exact sampling of the skew Gaussian process posterior avoids both (Benavoli et al., 2021c). How large the errors are at human noise levels is unmeasured (Section 27.4), so the switch matters most for very consistent answerers (inference).
Sources cited in Section 46.3 7
- Vessel et al. (2018) Stronger shared taste for natural aesthetic domains than for artifacts of human culture
- Liao et al. (2026) Efficient Human-in-the-Loop Optimization via Priors Learned from User Models
- Owaki et al. (2026) Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs
- Doumont et al. (2026) We Still Don't Understand High-Dimensional Bayesian Optimization
- Hvarfner et al. (2024) Vanilla Bayesian Optimization Performs Great in High Dimensions
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
46.4 Acquisition and query design #
The observation model settles what a query records and what happens when the person cannot answer (Chapter 20); Table 46.3 collects what the evidence says about choosing and presenting the queries.
| Recommendation | Basis | Depends on the GP and PBO framework? |
|---|---|---|
| Keep a fixed share of random or balanced queries; randomize left and right and positions in a gallery; rotate which option is the system's proposal | active learning is no better than random when preferences are unstable or the model class is wrong (Keswani et al., 2024); gaze has a causal effect on choice (Bhatnagar and Orquin, 2022); Glickman and Sharot (2025) | no |
| Monitor the connectivity and algebraic connectivity of the comparison graph over evaluated designs, and avoid isolated pairs | Shah et al. (2016); Hendrickx et al. (2019); Shao et al. (2026) (preprint) | partly: rank deficiency concerns the Laplace approximation, connectivity matters for any Bradley-Terry estimate |
| When users may differ, use rankings of three or more options instead of binary comparisons; penalize pairs that differ in many attributes at once | Chidambaram et al. (2026); Johnston et al. (2017) | no |
| Let people state goals and constraints in natural language, and use comparisons or rankings for fine judgments of appearance; make the state of the search visible and editable | Niwa et al. (2025); Peng et al. (2026) (preprint) | no |
Random and balanced queries. A simulation study of moral preference elicitation found that active learning rests on three premises (preferences that are stable and unaffected by the order of queries, a correct model class, and limited noise), and that when they fail, actively chosen queries do no better than random ones, or worse (Keswani et al., 2024). A fixed share of queries that the acquisition function does not choose costs some efficiency and buys a control group inside every session, free of the system's steering, against which the model's predictions can be checked (Section 46.8). Position matters too: steering how long people looked at an option raised its share of choices to 0.541 (95% confidence interval 0.523 to 0.560), a causal but small effect (Bhatnagar and Orquin, 2022), so randomizing positions and rotating which option the system proposes keeps such effects from adding up in one direction (inference).
The comparison graph. If evaluated designs and answered pairs fall apart into disconnected pieces, the answers say nothing about how the pieces compare, and the Laplace approximation becomes ill-conditioned (Section 18.5, Figure 27.1). The graph's algebraic connectivity, the second-smallest eigenvalue of its Laplacian matrix, is zero exactly then. Since EUBO's pairs tend to start new components (Shao et al., 2026) (Section 19.6), the check is cheap insurance: when the next pair would start a new component, swap one of its options for an already-compared design (inference).
Rankings and simple pairs. When users differ in ways the model does not know, binary comparisons cannot identify the user types, and rankings of at least three options can (Chidambaram et al., 2026). Guidance for stated-preference studies notes that respondents' consistency falls as the statistical efficiency and the complexity of the choice tasks rise (Johnston et al., 2017); in PBO, the complex tasks are pairs that differ in many attributes at once (inference).
Language for goals, comparisons for appearance. In the two studies of optimizer-led design ( and ; Section 45.3.3), steering through natural language gave lower workload than explicit constraints, and 90.9% of 187 requests in natural language described a desired outcome rather than a parameter (Niwa et al., 2025). For fine judgments of appearance, comparisons work better than words: in a 2026 preprint, personalization from just 8 pairwise judgments beat every baseline for 12 new users, including their own written preferences, with an aggregate win rate of 60.35% (Peng et al., 2026). The photo case study lets you try comparisons and line search on an image (Chapter 25).
Sources cited in Section 46.4 10
- Keswani et al. (2024) On the Pros and Cons of Active Learning for Moral Preference Elicitation
- Bhatnagar and Orquin (2022) A meta-analysis on the effect of visual attention on choice
- Glickman and Sharot (2025) How human–AI feedback loops alter human perceptual, emotional and social judgements
- Shah et al. (2016) Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
- Hendrickx et al. (2019) Graph Resistance and Learning from Pairwise Comparisons
- Shao et al. (2026) Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization
- Chidambaram et al. (2026) Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
- Johnston et al. (2017) Contemporary Guidance for Stated Preference Studies
- Niwa et al. (2025) Cooperative Design Optimization through Natural Language Interaction
- Peng et al. (2026) Efficient Personalization of Generative User Interfaces
46.5 Non-stationarity #
A preference that changes during a session breaks the model's central assumption. Two kinds of change need to be separated: drift, where the utility itself moves (through learning, adaptation, fatigue, or change caused by the queries), and a change in the choice rule, where the utility stays put but the way the person turns it into an answer changes, for instance from weighing all attributes to deciding on the single one that has come to matter (Section 37.5). Drift calls for a time-varying utility, a changed choice rule for a time-varying noise or link (inference).
| Recommendation | Basis | Depends on the GP and PBO framework? |
|---|---|---|
| Show the same pairs again at different lags, to separate test-retest noise from drift; when tuning a device, schedule a familiarization phase and down-weight early comparisons | Keswani et al. (2026); becoming an expert user took about 109 minutes of assisted walking (Poggensee and Collins, 2021) | no |
| Where needed, use a time kernel with a fast and a slow scale, or a state-space drift per parameter | inference, from adaptation at several time scales | the time kernel belongs to the GP framework; state-space drift does not |
Repeated pairs. When the same moral pairwise questions were asked in three to five sessions, 6% to 20% of answers flipped, and for some participants the fitted decision model itself changed (Keswani et al., 2026). One repeated pair cannot tell noise from change; pairs repeated at several lags can, because noise does not depend on the lag and drift does (Figure 46.1).
Things to try:
- At the defaults the curve falls from 72% at a lag of 2 answers to 61% at a lag of 40, because the utility differences drift; with 30 people and 4 repeated pairs per lag, the measured points show the decline. Set drift to 0, and the curve is flat at 73%: noise makes people disagree with themselves equally at every lag.
- Raise Incomplete (the share of incomplete answers) to 0.4 with drift at 0. The curve stays flat and drops to 58%: answers with no preference behind them look like extra noise unless "can't compare" is offered.
- Return Incomplete to 0, keep drift at 0, and set Induced (induced change) to 1. The curve falls again, from 92% to 74%, because at short lags the person repeats a choice that the first choice made more attractive. A falling curve does not distinguish drift from induced change; only comparing query orders does (Section 45.4.1).
- Reset the figure, then set People to 5 and Repeats (repeated pairs per lag and person) to 1. The intervals span about ±33 percentage points. A diagnostic built on repeated pairs needs on the order of a hundred repeated pairs per lag, pooled over people (inference).
Familiarization. People adapt to a device over longer than a typical tuning session (in exoskeleton walking, training contributed about half of a 39% reduction in metabolic cost (Poggensee and Collins, 2021)), and comparisons made before they have adapted measure a different person (Section 24.4).
Models of change. Where drift is real, a time kernel with a fast and a slow scale, or a state-space model in which each parameter drifts, can represent it (inference); over months, a stable component plus deviations that follow events and then decay fits the evidence better than a random walk (Chapter 42). Drifting utilities are not yet modeled in PBO (Section 27.2), and the theory that exists is for finite arms (Section 29.10).
Sources cited in Section 46.5 2
- Keswani et al. (2026) Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
- Poggensee and Collins (2021) How adaptation, training, and customization contribute to benefits from exoskeleton assistance
46.6 Stopping #
When should a session end? The tempting answer is when the posterior has concentrated: the model is confident about the best design, and EUBO, asked to choose a pair, would show the incumbent against itself (Exercise 19.2). That is not enough, because the posterior concentrates whenever the answers are consistent, and choice-induced revaluation makes answers consistent without any stable preference behind them (Figure 45.2, Exercise 45.2).
| Recommendation | Basis | Depends on the GP and PBO framework? |
|---|---|---|
| Do not treat posterior concentration as evidence that the preference has stabilized, since revaluation after choosing also concentrates it; stop on a plateau in repeated-pair agreement, a delayed retest that reproduces the result, and an endorsement check | inference, from Zylberberg et al. (2024) | no |
| Carry over stopping rules based on the probability of regret, or on cost, only after testing them with a pairwise likelihood and a person's changing cost | Wilson (2024); Xie et al. (2026) | yes |
Rules from scalar Bayesian optimization. Scalar rules stop when a design is -optimal with probability at least (Wilson, 2024), or when no point's expected improvement exceeds the cost of evaluating it (Xie et al., 2026); both assume a correctly specified Gaussian process prior (Section 30.7, Section 14.7). Preference-based reward learning has an optimal stopping rule for parametric reward models (Bıyık et al., 2019). None has been carried over to PBO's pairwise Gaussian process likelihood, and a person's cost per query is not constant: fatigue raises it, and learning the task lowers it. Until a rule is carried over, Algorithm 46.3 is the practical one.
Response times as a stopping signal. Answers usually get faster late in a session. That can mean the preference has settled, but people also stop deliberating when the expected gain in confidence is no longer worth the effort (Bénon et al., 2024), so faster answers need the same caution as posterior concentration (inference).
People stop on their own. In a three-month field deployment, 415 of 549 evaluation sequences stopped at the first iteration, and only 16 of the remaining 134 reached a satisfying result (Ou et al., 2022). A person's "good enough" and a model's converged posterior are different signals, and a stopping rule designed for the model must also survive the person's own decision to quit.
- Continue while the model's own criterion says more answers are worth having, for example while a challenger still has a real chance of beating the incumbent.
- Once it is met, check that agreement on repeated pairs has stopped rising over the last few repeats.
- Retest the incumbent, without the system's framing, against two or three designs the person rejected earlier. If it loses, continue.
- Ask the person whether they endorse the result. Economic experiments find that people revise their choices when asked to reflect on principles they themselves endorse (Nielsen and Rehbeck, 2022), so the question is informative, though not by itself a justification (Section 45.3.3).
- Stop, and schedule the delayed retest of Algorithm 46.2.
Sources cited in Section 46.6 7
- Zylberberg et al. (2024) Value construction through sequential sampling explains serial dependencies in decision making
- Wilson (2024) Stopping Bayesian Optimization with Probabilistic Regret Bounds
- Xie et al. (2026) Cost-aware Stopping for Bayesian Optimization
- Bıyık et al. (2019) Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
- Bénon et al. (2024) The online metacognitive control of decisions
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
- Nielsen and Rehbeck (2022) When Choices Are Mistakes
46.7 Evaluation #
How should a new PBO method or system be evaluated? Current practice is weak, with scalar test functions, incompatible noise models and definitions of regret, and no shared benchmark (Section 45.1.4, Section 31.4). Three recommendations, all independent of the framework, address it, and Table 46.6 collects them with the others into a reporting checklist.
Strong baselines. Most positive results for interactive Bayesian optimization in design come from comparisons with sliders, parameter panels, or random queries; against skilled manual work or a similar optimizer, final quality usually does not differ, and these systems are best summarized as reaching the same result at lower cost (Section 32.8). A method that has not been compared with manual tuning, random search, and a linear model cannot claim to be needed.
Acceleration factors. Self-driving laboratories report the acceleration factor, how many experiments a strategy needs to reach a target relative to a reference strategy: across 42 studies and 63 benchmarks, the median reported factor was 6, with a range of 1.3 to 100 (Adesiji et al., 2026). PBO studies with people can report the same number relative to manual or random search, which makes results comparable across tasks (inference).
Simulated users. Test the assumptions of a simulated user on at least a few real people. Simulated agents agreed with the people they stood in for on only about 50% of choices (Schoinas et al., 2025) (Figure 45.1), and language-model simulations of users recover population-level rankings almost perfectly while agreeing with individuals at a Kendall of only 0.11 to 0.22, against 0.57 for the agreement between people's own ratings and their own rankings, in a 2026 preprint (Kirk et al., 2026). A result obtained with a simulated user is a result about that simulation until it has been checked on people (Section 31.7).
| Item | What to report | Why |
|---|---|---|
| Baselines | manual tuning, random or coarse grid search, a linear utility model | positive results mostly come from weak comparisons (Section 32.8) |
| Setting | dimension, noise model, definition of regret, software and version, number of repetitions | rankings of methods flip with these (Section 45.1.1) |
| Decision maker | a person or a simulated user; what preference information was collected; how stopping was decided | the norms of interactive multi-objective optimization (Section 40.11) |
| Cost | an acceleration factor relative to manual or random search, in comparisons and in minutes | the typical benefit is lower cost, not a better result |
| Noise audit | agreement on repeated pairs at several lags; the false-preference rate on placebo pairs | separates noise from drift (Figure 46.1) |
| Influence | the shift in preference during the session, against a group with random or balanced query order | the output is an estimate and an intervention (Section 45.5) |
| Delayed outcome | an unframed retest a week later, and, where possible, the result in actual use | a choice in the session is not the same as satisfaction in use (Section 37.2) |
| Data | individual-level comparisons with timestamps, presentation order, response times, and repeated pairs | no such public data set exists (Section 31.5) |
Sources cited in Section 46.7 3
- Adesiji et al. (2026) Benchmarking self-driving labs
- Schoinas et al. (2025) Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants
- Kirk et al. (2026) PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
46.8 Ethics and legitimacy #
A system that learns a preference while shaping it needs rules about how much shaping is acceptable. Section 45.5.2 set out the candidate conditions; Table 46.7 turns them into practice.
| Recommendation | Basis | Depends on the GP and PBO framework? |
|---|---|---|
| Report the shift in preference within the session as well as regret, with a control group that receives a random or balanced query order | Dean and Morgenstern (2022); Carroll et al. (2022) | no |
| Audit whether the objective and the stopping rule reward the user for becoming predictable; label proposals, show the history accurately, and allow a reset | Carroll et al. (2023); Williams et al. (2025) | no |
| Before the session, record the person's attitude to being changed; after it, retest after a delay and without framing; for morally weighty decisions, map the options and their trade-offs rather than output a decision | Pettigrew (2023); Kanwal and Tran (2026) (workshop paper); Chapter 41 | no |
Measure the shift, not only the regret. Because low regret can be achieved trivially by moving the user's preferences (Section 45.2), regret alone cannot certify a system that can influence its user. Measuring how far preferences moved during the session, against a control group whose queries were random or balanced, makes the influence visible, and the shift can then be bounded (Carroll et al., 2022).
Audit the incentives. Since learners optimized on user feedback learn to target the users easiest to influence (Section 45.5.1), a team can ask whether its acquisition function and stopping rule reward answers becoming more predictable; labeling proposals, showing the history accurately, and allowing a reset keep the influence in view. A system that chooses which candidates to show and also benefits from particular outcomes is in the position of a sender in the economics of persuasion, whose preferences differ from the user's (Section 40.6).
Before and after. Recording the person's attitude before the session and retesting after a delay without the system gives both the earlier and the changed self a say (inference; Section 45.5.2). For decisions with moral weight, the evidence on autonomy argues for a system that lays out the options and trade-offs and leaves the decision to the person (Section 41.2).
For systems deployed to consumers, these steps also serve compliance. Article 25 of the EU Digital Services Act prohibits providers of online platforms from designing interfaces that materially distort or impair users' ability to make free and informed decisions (European Union, 2022b). An acquisition function that systematically reinforces the current favorite might fall within that wording, and logging a balanced query order and ending with an explicit endorsement check serve scientific validity and that requirement at once (inference; this is our reading of the text, not legal advice; Section 42.2.3).
Sources cited in Section 46.8 7
- Dean and Morgenstern (2022) Preference Dynamics Under Personalized Recommendations
- Carroll et al. (2022) Estimating and Penalizing Induced Preference Shifts in Recommender Systems
- Carroll et al. (2023) Characterizing Manipulation from AI Systems
- Williams et al. (2025) On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
- Pettigrew (2023) Nudging for changing selves
- Kanwal and Tran (2026) Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
- European Union (2022b) Regulation (EU) 2022/2065 on a Single Market for Digital Services (Digital Services Act)
46.9 When a simpler method should win #
Random search learns nothing from its own results, and in several settings nothing beats it by much (Chapter 1). Table 46.8 lists the situations in which a simpler method should be the default or a required baseline, and whether the advice comes from the problem itself, from direct evidence, or from inside the Gaussian process and PBO framework.
| Situation | Default or baseline | Why | Source of the advice |
|---|---|---|---|
| Few parameters (about 4 to 6), each evaluation must be felt with the body, the optimum may be flat | manual self-tuning | self-tuning with a thumbstick reduced metabolic cost by 16.6%, relative to walking with the exoskeleton in a zero-torque condition, in about 11 minutes (Schäfer et al., 2026); self-tuning of ankle torque converged in 105 seconds within a trial (Ingraham et al., 2022); for 2 of 3 amputees, the self-selected setting differed between trials on the same day (Díaz et al., 2026) | fit to the problem, and evidence |
| A flat optimum, candidates hard to tell apart | a coarse grid or discrete candidates | three users of a discrete version of PBO for a prosthesis recognized their final setting in 93% of validation trials, the one user of a continuous version in 67% (Taddei et al., 2026) (preprint) | fit to the problem |
| Attributes can be stated explicitly, utility roughly linear and additive | a linear utility model or adaptive conjoint analysis | adaptive Bayesian question selection with response error is a mature method (Sauré and Vielma, 2019), with a documented estimation bias for utility-balanced questions (Hauser and Toubia, 2005); a linear model reached the state of the art on high-dimensional benchmarks (Doumont et al., 2026) | fit to the problem; whether a linear model suffices for human pairwise data is untested |
| Nominal dimension 20 or more | a low-dimensional representation (a generative latent space, a latent space that respects feasibility) | a 5-dimensional feasibility-aware latent space gave higher shape similarity and more feasible suggestions than 9 raw parameters (Owaki et al., 2026) (preprint) | fit to the problem |
| Several measurable objectives, with anchoring and loss framing to avoid | an interactive multi-objective method, for example NAUTILUS, which starts from a dominated solution and improves every objective at each step | Miettinen et al. (2010); anchoring is stronger when deciding on someone else's behalf (Halstead et al., 2026) | fit to the problem |
| Collecting preference data with a strong pretrained prior, or with very noisy data | random on-policy sampling | in online direct preference optimization of language models, active selection improved the proxy win rate only negligibly over random selection (Oh et al., 2026b) (workshop paper); with noisy data, random search with large batches can be a good choice (Siivola et al., 2021) | evidence |
| Ideation, when goals have not yet formed | tools for divergence, or archives of diverse solutions | an AI image generator during ideation led to more fixation and fewer, less diverse, and less original ideas () (Wadinambiarachchi et al., 2024); quality-diversity search with human feedback returns diverse, high-quality archives (Ding et al., 2024) | fit to the problem |
| The goal is a conclusion about a population | a design that mixes in random queries | for population parameters, adaptive Bayesian designs were consistently less precise than random designs (Gibbard and Sadlier, 2025) | evidence |
| Only perceptual comparison possible, no numerical objective, low dimension, tens to about 200 comparisons, manual search infeasible (for example, the parameters cannot be manipulated directly) | PBO with qEUBO and PairwiseGP |
the best-maintained implementation; positive human studies exist in this range, but the controls were mostly sliders or random queries | the choice of implementation comes from the framework and the software ecosystem |
The case studies work several of these situations end to end: the exoskeleton case re-enacts the first row (Section 24.5), the classifier case shows random search holding its own against Bayesian optimization (Section 22.3), and the photo case (Chapter 25) is an instance of the last row.
The overall verdict, which comes from the evidence and not from the framework, is that in studies with people, an advantage of PBO over strong simple baselines has not been shown directly. New studies should therefore make manual tuning, random or coarse grid search, and a linear utility model required controls; comparing only with sliders or with another PBO variant cannot answer whether PBO is needed at all.
Sources cited in Section 46.9 15
- Schäfer et al. (2026) User preference-based human-in-the-loop tuning of exoskeleton assistance during walking
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Díaz et al. (2026) User preference in the personalized control of an ankle prosthesis: a case study
- Taddei et al. (2026) Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis
- Sauré and Vielma (2019) Ellipsoidal Methods for Adaptive Choice-Based Conjoint Analysis
- Hauser and Toubia (2005) The Impact of Utility Balance and Endogeneity in Conjoint Analysis
- Doumont et al. (2026) We Still Don't Understand High-Dimensional Bayesian Optimization
- Owaki et al. (2026) Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs
- Miettinen et al. (2010) NAUTILUS method: An interactive technique in multiobjective optimization based on the nadir point
- Halstead et al. (2026) Multiobjective Optimisation for Others: How Anchoring Effects Change Based on Who Guides the Interaction
- Oh et al. (2026b) Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
- Siivola et al. (2021) Preferential Batch Bayesian Optimization
- Wadinambiarachchi et al. (2024) The Effects of Generative AI on Design Fixation and Divergent Thinking
- Ding et al. (2024) Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization
- Gibbard and Sadlier (2025) Optimal adaptive Bayesian design in choice experiments: performance for the population versus individuals
46.10 Settled, contested, missing #
Settled. In exoskeleton tuning with few parameters, self-tuning reaches metabolic savings of the same order as algorithmic tuning. For population-level parameters, adaptive Bayesian designs are less precise than random ones. A forced choice cannot distinguish indifference, incompleteness, and noise. A disconnected comparison graph leaves relative utilities between its components undetermined.
Contested. Whether random selection is hard to beat outside the online preference optimization of language models (the main evidence is a workshop paper). Whether the Laplace approximation's errors matter at human noise levels. Whether dimension-scaled lengthscale priors repair the default model under pairwise feedback.
Missing. A test of any of these recommendations as a whole, in particular of the instrumented session and the stopping rule. A study in which PBO beats manual tuning, random search, and a linear utility model on a preregistered endpoint. A measured share of "can't compare" answers in design tasks. A public data set with which these recommendations could be checked offline.
46.11 Exercises #
A team wants to personalize 12 parameters of a hearing aid's sound processing for each user, in sessions of about 40 comparisons spread over two clinic visits. Work through Algorithm 46.1. What would you build, what baselines would you require, and what would you add to the session?
Solution
Question 1. Measurable quantities such as speech intelligibility do not capture what the user prefers to hear; they can serve as constraints, and the comparisons carry the preference. Question 2. Twelve exposed parameters exceed the guideline of about 10, and 40 comparisons carry at most 40 bits, so reduce what the user faces first, for example to a few perceptual directions built from earlier users' data, or to a short list of discrete presets. Question 3. With a few directions, a person adjusting sliders may do as well; self-adjustment is the first baseline, alongside random presets and a linear utility model. If PBO is still the choice, instrument it as in Algorithm 46.2; the two visits make repeated pairs and an unframed retest across visits cheap, and they also capture adaptation to the device (Section 46.5).
Six designs have been compared in the pairs (1, 2), (3, 4), (5, 6), and (1, 3). How many connected components does the comparison graph have, what is its algebraic connectivity, and which kind of query would you ask next? Why does this matter less for a model with a strong prior?
Solution
The edges join {1, 2, 3, 4} into one component, and {5, 6} form a second, so there are two components and the algebraic connectivity, the second-smallest eigenvalue of the graph Laplacian, is 0. No answer so far says how 5 and 6 compare with the other four; any query that pairs 5 or 6 with one of 1 to 4 connects the graph. A Gaussian process prior with a long lengthscale links utilities of nearby designs even without comparisons, so the posterior is still well defined, but the offset between the two groups is then set by the prior rather than by the person (Section 27.6).
After 25 comparisons the posterior has concentrated on one design. Repeated pairs agree 90% of the time at a lag of 2 answers and 60% at a lag of 20. What do you conclude, and what do you do next?
Solution
Agreement that falls with the lag means the answers are changing: either the utility is drifting, or earlier choices made some options temporarily more attractive (Figure 46.1). Either way the concentrated posterior may reflect a moving target, or the system's own influence, rather than a stable preference, so do not stop on it (Section 46.6). Check the answers to the random share of queries against the model's predictions, retest the incumbent without framing against designs rejected early, and schedule a delayed retest; in a study, only a comparison with a group that received a random query order tells drift from influence.
Further reading #
- Schäfer et al. (2026) is the clearest demonstration that a strong simple baseline can match algorithmic tuning with a person in the loop.
- Keswani et al. (2024) states the premises on which active preference elicitation depends and shows what happens when they fail.
- O'Mahony and Wichchukit (2017) reviews how sensory science learned to handle forced choice, "no preference" answers, and placebo pairs, decades before PBO.
- Dean and Morgenstern (2022) and Carroll et al. (2022) show why regret is not enough when the system can move the preference it is learning.
- Adesiji et al. (2026) sets out the acceleration and enhancement factors that make results comparable across tasks.
References
- (2026). Benchmarking self-driving labs. Digital Discovery. Cited in §46.7
- (2024). The online metacognitive control of decisions. Communications Psychology. doi:10.1038/s44271-024-00071-y. Cited in §46.6
- (2022). A meta-analysis on the effect of visual attention on choice. Journal of Experimental Psychology: General. Cited in §46.4
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Cited in §46.6
- (2024). Meta-analysis of Empirical Estimates of Loss Aversion. Journal of Economic Literature. Cited in §46.2
- (2022). Estimating and Penalizing Induced Preference Shifts in Recommender Systems. ICML 2022. Cited in §46.8
- (2023). Characterizing Manipulation from AI Systems. Equity and Access in Algorithms, Mechanisms, and Optimization. Cited in §46.8
- (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Cited in §46.4
- (2022). Preference Dynamics Under Personalized Recommendations. EC 2022. Cited in §46.8
- (2026). User preference in the personalized control of an ankle prosthesis: a case study. Journal of NeuroEngineering and Rehabilitation. doi:10.1186/s12984-026-01931-w. Cited in §46.9
- (2024). Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization. ICML 2024. Cited in §46.9
- (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). Cited in §46.3 §46.9
- (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Cited in §46.2
- (2022b). Regulation (EU) 2022/2065 on a Single Market for Digital Services (Digital Services Act). Official Journal of the European Union (EUR-Lex). non-peer-reviewed Cited in §46.8
- (2023). The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types. AAAI. Cited in §46.2
- (2025). Optimal adaptive Bayesian design in choice experiments: performance for the population versus individuals. Marketing Letters. Cited in §46.9
- (2025). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour. doi:10.1038/s41562-024-02077-2. Cited in §46.4
- (2026). Multiobjective Optimisation for Others: How Anchoring Effects Change Based on Who Guides the Interaction. Journal of Multi-Criteria Decision Analysis. Cited in §46.9
- (2005). The Impact of Utility Balance and Endogeneity in Conjoint Analysis. Marketing Science. Cited in §46.9
- (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. Cited in §46.4
- (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. Cited in §46.3
- (2022). The role of user preference in the customized control of robotic exoskeletons. Science Robotics. Cited in §46.9
- (2017). Contemporary Guidance for Stated Preference Studies. Journal of the Association of Environmental and Resource Economists. Cited in §46.4
- (2026). Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction. AAAI-26 Workshop on Machine Ethics. workshop paper Cited in §46.8
- (2024). Stochastic Monotonicity and Random Utility Models: The Good and The Ugly. arXiv preprint 2409.00704. preprint Cited in §46.2
- (2024). On the Pros and Cons of Active Learning for Moral Preference Elicitation. AIES. Cited in §46.4
- (2026). Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback. AAAI. Cited in §46.5
- (2026). PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users. arXiv. preprint Cited in §46.7
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Cited in §46.1
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Cited in §46.3
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Cited in §46.1
- (2010). NAUTILUS method: An interactive technique in multiobjective optimization based on the nadir point. European Journal of Operational Research. Cited in §46.9
- (2022). When Choices Are Mistakes. American Economic Review. Cited in §46.6
- (2026). Revealed Incomplete Preferences. Working paper (author's website). working paper Cited in §46.2
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Cited in §46.4
- (2017). The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection. Trends in Food Science & Technology. Cited in §46.2
- (2026b). Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs. ICLR 2026 Workshop: I Can't Believe It's Not Better (ICBINB). workshop paper Cited in §46.9
- (2022). Indifference, indecisiveness, experimentation, and stochastic choice. Theoretical Economics. Cited in §46.2
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §46.6
- (2026). Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs. arXiv. preprint Cited in §46.3 §46.9
- (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Cited in §46.4
- (2023). Nudging for changing selves. Synthese. Cited in §46.8
- (2021). How adaptation, training, and customization contribute to benefits from exoskeleton assistance. Science Robotics. Cited in §46.5
- (2019). Ellipsoidal Methods for Adaptive Choice-Based Conjoint Analysis. Operations Research. Cited in §46.9
- (2025). Preference Learning with Response Time: Robust Losses and Guarantees. NeurIPS. Cited in §46.2
- (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Cited in §46.9
- (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Cited in §46.7
- (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Cited in §46.4
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Cited in §46.4
- (2025b). Early versus late noise differentially enhances or degrades context-dependent choice. Nature Communications. doi:10.1038/s41467-025-59140-3. Cited in §46.2
- (2022). High-value decisions are fast and accurate, inconsistent with diminishing value sensitivity. Proceedings of the National Academy of Sciences. Cited in §46.2
- (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Cited in §46.2
- (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Cited in §46.9
- (2026). Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis. arXiv. preprint Cited in §46.9
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §46.3
- (2018). Stronger shared taste for natural aesthetic domains than for artifacts of human culture. Cognition. Cited in §46.3
- (2024). The Effects of Generative AI on Design Fixation and Divergent Thinking. Proceedings of the CHI Conference on Human Factors in Computing Systems. Cited in §46.9
- (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Cited in §46.8
- (2024). Stopping Bayesian Optimization with Probabilistic Regret Bounds. NeurIPS 2024. Cited in §46.6
- (2026). Cost-aware Stopping for Bayesian Optimization. International Conference on Machine Learning. Cited in §46.6
- (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Cited in §46.2 §46.6