Bayesian Optimization
Part VII: People in the Loop
中文

Built Environments, Science, and Industry

Chapter 32 and Chapter 33 followed preferential optimization where people were in the loop by necessity: a design only its user can judge, a device only its wearer can feel. This chapter covers the remaining applications, where the case for asking a person is less obvious and the evidence takes a different shape. In buildings and vehicles, comfort and driving style are personal, but the studies mostly use simulated occupants and drivers. In science and engineering, preferences inject an expert's judgment into an optimization that otherwise has a measured objective, and the experts are usually one to four people. In industry, the software is ready, and the public record consists of benchmarks and motivating statements. The chapter then compares the human evidence across every application domain of this part, and ends by checking claims about these applications that circulate in secondary accounts.

34.1 Buildings and vehicles #

34.1.1 Thermal comfort and daylight #

Comfort differs between people, and occupants answer comparisons about it easily, which makes buildings a natural target. The evidence is almost entirely simulated. A Purdue University group learned individual visual satisfaction with office daylighting from comparative preferences between 2017 and 2020 (Xiong et al., 2017; Xiong et al., 2018; Xiong et al., 2020); we could not obtain the sample sizes or the number of queries. Awalgaonkar et al. (2019), a preprint, treated an occupant's answer ("I would like it warmer", "cooler", or "I am satisfied") as the sign of the derivative of their utility, used a Gaussian process constrained to a single peak, and validated it with synthetic occupants.

POP-BO, an optimistic method of preferential Bayesian optimization (PBO) with a bound on cumulative regret (Xu et al., 2024b), includes a thermal comfort task in simulation, with preferences generated from the predicted mean vote (PMV), a standard formula for a group's average comfort rating, over 2 or 4 variables such as air temperature and air speed. qEUBO (EUBO for queries of qq options, Section 19.4) found a slightly better final solution, but its cumulative regret was almost twice that of POP-BO, and the authors argue that online control of heating, ventilation, and air conditioning (HVAC) should care about cumulative regret, the total discomfort over the whole run (Chapter 13). Three more recent preprints are simulations too: contextual PBO raised occupant utility by up to 23% over two simulated months on BOPTEST, a building simulation platform (Wang et al., 2025d); a real-time preference-optimizing controller was evaluated on simulated thermal comfort (Wang et al., 2025c); and a consensus Bayesian optimization robust to social influence uses thermal comfort as one application (Adachi et al., 2025).

Every building study of PBO from 2024 to 2026 used simulated occupants. That is a gap, not a negative result: the next useful study is a trial of preference-based HVAC control with real occupants, reporting cumulative discomfort as well as the final setting (inference).

34.1.2 Vehicles and controller calibration #

People misjudge their own driving. In Basu et al. (2017), not an optimization study, users preferred a driving style clearly more defensive than how they drove, preferred the style they believed was their own over their actual one, and preferred different styles in different scenarios, so a system that learns from how people drive and one that learns from what they say they prefer reach different places.

One study with real drivers. Ran et al. (2023) personalized lane-centering trajectories online from preferences, modeling the uncertainty in each answer; in a fixed-base driving simulator, 29 drivers (15 inexperienced, 14 experienced) converged after 11.1 ± 4.6 queries on average, almost all in fewer than 20, and the estimated utility agreed with their own rating scores.

Controller calibration in simulation. A model predictive controller (MPC) chooses each action by optimizing over a short prediction of the future, with weights an engineer usually calibrates by trial and error. Zhu et al. (2021) calibrated MPC controllers with GLISp, the radial basis function method of Section 33.6, on a simulated reactor and a simulated lane-keeping car (50 experiments each); for the car, the authors could not write a scoring function usable for automatic calibration even after much trial and error. In the later studies the decision maker was synthetic or an author. C-GLISp adds feasibility labels for unknown constraints, with a synthetic decision maker in its benchmarks and the first author as calibrator in its driving case (Zhu et al., 2022); Theiner et al. (2025) guided PBO with a virtual decision maker built from real driving data, again with a synthetic main decision maker, and converged faster than standard PBO, which needed about 70 iterations, with a multi-fidelity follow-up (Theiner et al., 2026); and two preprints tuned a vehicle suspension with a synthetic user (Cercola et al., 2026b) and the slip control of an electric race car, with simulation results only (de Vries et al., 2024). The in-car interface studies optimized from ratings are in Section 32.1.4.

Sources cited in Section 34.1 16
  1. Xiong et al. (2017) Personalized visual satisfaction profiles from comparative preferences using Bayesian inference
  2. Xiong et al. (2018) Inferring personalized visual satisfaction profiles in daylit offices from comparative preferences using a Bayesian approach
  3. Xiong et al. (2020) Efficient learning of personalized visual preferences in daylit offices: An online elicitation framework
  4. Awalgaonkar et al. (2019) Learning Personalized Thermal Preferences via Bayesian Active Learning with Unimodality Constraints
  5. Xu et al. (2024b) Principled Preferential Bayesian Optimization
  6. Wang et al. (2025d) Personalized Building Climate Control with Contextual Preferential Bayesian Optimization
  7. Wang et al. (2025c) Human-in-the-loop: Real-time Preference Optimization
  8. Adachi et al. (2025) Bayesian Optimization for Building Social-Influence-Free Consensus
  9. Basu et al. (2017) Do You Want Your Autonomous Car To Drive Like You?
  10. Ran et al. (2023) Online Personalized Preference Learning Method Based on In-Formative Query for Lane Centering Control Trajectory
  11. Zhu et al. (2021) Preference-based MPC calibration
  12. Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
  13. Theiner et al. (2025) Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles
  14. Theiner et al. (2026) Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models
  15. Cercola et al. (2026b) Regularized GLISp for sensor-guided human-in-the-loop optimization
  16. de Vries et al. (2024) A Human-optimized Model Predictive Control Scheme and Extremum Seeking Parameter Estimator for Slip Control of Electric Race Cars

34.2 Science and engineering with expert preferences #

In the sciences, preferences mostly bring an expert's judgment into a Bayesian optimization whose objective would otherwise be a measured scalar, and in each study the real experts number one to four. Section 15.3 covers Bayesian optimization in experimental science with measured objectives, and Chapter 23 works through a chemical reaction end to end.

34.2.1 Materials and physical sciences #

Expertise visibly changes the outcome. The projective preferential Bayesian optimization of Mikkola et al. (2020) asks the user to pick the best point along a one-dimensional projection with a slider (Section 20.3). It was used to find the position and orientation of a camphor molecule on a copper surface, Cu(111), with each candidate's energy checked by density functional theory, a quantum-mechanical calculation of the energy of a molecular configuration. Every user answered 24 queries. The structures preferred by two materials-science experts relaxed to energy minima of about -1.007 to -1.030 eV, while those of a non-expert relaxed to shallower local minima of about -0.762 to -0.771 eV. The authors describe this as a clear split by expertise, and the method could also tell a person's choices from those of a random robot.

Votes on spectra. BOARS turns an operator's up or down votes on measured spectra into an objective by Thurstone-Mosteller scaling (Section 16.3), with the human leading early and the algorithm later, demonstrated on a live atomic force microscope (Biswas et al., 2024). A 2026 preprint, px-BO, fits a Bradley-Terry model (Section 16.4) to the votes and then lets a surrogate vote in the human's place, with periodic human checks (Biswas et al., 2026).

Explanations, and the risk of over-trust. In CoExBO (Section 32.7), an expert chooses between two candidates shown with explanations (Adachi et al., 2024). On battery electrolyte design with 4 participants, the explanations improved the experts' accuracy; the authors note that the gains assume the expert's comparative knowledge is accurate, and that experts tended to expect the surrogate to understand the problem like an oracle, a form of over-trust.

Simulated experts. BOAP models an expert's pairwise preferences over "abstract properties" that cannot be measured, for lithium-ion electrode manufacturing, but the preferences were simulated from published data sets (Arun Kumar A V et al., 2024). In the multi-objective optimization of scanning probe microscopy by Liu and Kalinin (2025), the human steers by objective weights and reference points, not comparisons. Deneault et al. (2025) used PBO to optimize subjective qualities of 3-D prints, judged by a person rather than measured by sensors.

34.2.2 Chemistry and drug discovery #

Chemistry holds the only large data set of expert preferences in this chapter. MolSkill collected more than 5,000 pairwise comparisons of molecules from 35 chemists at Novartis (in wet-lab, computational, and analytical roles) over several months, choosing each round by active learning (Choung et al., 2023). Agreement was moderate. In preliminary rounds, the inter-rater agreement on 200 pairs, measured by Fleiss' kappa, a chance-corrected agreement among several raters where 0 is chance and 1 is perfect, was 0.40 and 0.32; the intra-rater agreement on 20 repeated pairs, measured by Cohen's kappa, was 0.60 and 0.59, and for some chemists as low as 0.16 to 0.27. The area under the ROC curve for classifying pairs, the probability that the model ranks a random preferred molecule above a random non-preferred one, rose from about 0.6 with 1,000 pairs to above 0.74 with 5,000, without reaching a plateau. The learned score captured drug-likeness better than QED, the quantitative estimate of drug-likeness chemists commonly use. MolSkill is a neural network that learns to rank, not Bayesian optimization.

The other chemistry studies use simulated experts, a single expert, or none. Sundin et al. (2022) learned a chemist's scoring function actively, improving significantly in fewer than 200 queries in two cases with simulated ground truth, with one medicinal chemist in a demonstration who judged one molecule at a time, not pairs. CheapVS, a workshop paper, guided screening of 100,000 molecules with chemists' pairwise preferences over trade-offs among properties: screening 6% of the library recovered 16 of the 37 known drugs for one target (EGFR) and 37 of the 58 for another (DRD2), but with a synthetic utility function, preliminary human data (expert rankings for EGFR about 80% accurate), and a number of experts we could not find (Dang et al., 2025). Kristiadi et al. (2024b), a workshop paper, simulated experts' pairwise labels with a scoring function; the "preferential Bayesian optimization" for protein design of Hawkins-Hooker et al. (2023), also a workshop paper, takes its comparisons from measured fitness; and Haltia et al. (2026), a preprint, choose between an expensive evaluation and a query to an expert by their costs.

34.2.3 Dividing the work between people and algorithms #

Kanarik et al. (2023) compared people and Bayesian optimization in a virtual process game for designing a semiconductor plasma etching process. Engineers did well early, the algorithm was far more cost-efficient close to tight tolerances, and a "human first-computer last" strategy halved the cost of reaching the target compared with engineers alone. The objective was a measured scalar, so how far this carries over to preference-driven design is uncertain. And Weichert et al. (2025), a preprint, document an industrial case in which adding expert knowledge made the optimization fail.

What it means. Agreement of about 0.3 to 0.4 between chemists and about 0.6 within one says that experts' pairwise intuition can be learned but is only moderately shared, which favors models per expert or mixture models over pooled data (inference; Section 27.2). Mikkola et al.'s split between experts and a non-expert, and CoExBO's assumption that expert knowledge is accurate, point to an untested failure mode: the expert who is wrong or overconfident. CoExBO's no-harm guarantee is the main design response so far (inference), and a method that relies on expert preferences could be tested with at least one deliberately misinformed expert (inference).

Sources cited in Section 34.2 15
  1. Mikkola et al. (2020) Projective Preferential Bayesian Optimization
  2. Biswas et al. (2024) A dynamic Bayesian optimized active recommender system for curiosity-driven partially Human-in-the-loop automated experiments
  3. Biswas et al. (2026) Human-AI Collaborative Autonomous Experimentation With Proxy Modeling for Comparative Observation
  4. Adachi et al. (2024) Looping in the Human Collaborative and Explainable Bayesian Optimization
  5. Arun Kumar A V et al. (2024) Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties
  6. Liu and Kalinin (2025) Pareto-Optimal Experimentation: Human-Guided Multi-Objective Bayesian Optimization in Scanning Probe Microscopy
  7. Deneault et al. (2025) Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities
  8. Choung et al. (2023) Extracting medicinal chemistry intuition via preference machine learning
  9. Sundin et al. (2022) Human-in-the-loop assisted de novo molecular design
  10. Dang et al. (2025) Preferential Multi-Objective Bayesian Optimization for Drug Discovery
  11. Kristiadi et al. (2024b) How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization?
  12. Hawkins-Hooker et al. (2023) Preferential Bayesian Optimisation for Protein Design with Ranking-Based Fitness Predictors
  13. Haltia et al. (2026) Elicitation-Augmented Bayesian Optimization
  14. Kanarik et al. (2023) Human–machine collaboration for improving semiconductor process development
  15. Weichert et al. (2025) When Less is More: A Story of Failing Bayesian Optimization Due to Additional Expert Knowledge

34.3 Industry #

The main industrial line: preference exploration at Meta. BOPE (preference exploration for Bayesian optimization with multiple outcomes) alternates between running experiments and asking a decision maker to compare predicted outcome vectors (Lin et al., 2022). The paper is motivated by A/B testing workflows, but its experiments use a vehicle safety problem (5 inputs, 3 outputs), a car cab design problem (7 inputs, 9 outputs), and the DTLZ2 and OSY test problems, with a noisy synthetic utility standing in for a human decision maker. qEUBO extended the expected utility of the best option to queries of qq options (Astudillo et al., 2023). Both are in BoTorch (Meta Platforms, Inc., 2026d), and between January and June 2026 Ax versions 1.2 to 1.3 added preference optimization configuration, BOPE utility tracking, qEUBO, and pairwise Gaussian processes (Meta Platforms, Inc., 2026l) (Section 31.1.2). LILO, which turns natural-language feedback into preference signals with a language model, was evaluated with a language model standing in for the decision maker (Kobalczyk et al., 2026) (Chapter 35). Other industry-facing applications include banner ads (Section 32.1.1), RankTuner for electronic design automation tools (Xu et al., 2026), contextual dueling bandits for recommendation (Sankagiri et al., 2026), multi-criteria decision support (Huber et al., 2025), and hearing-aid presets (Vyas et al., 2022).

Public evidence of deployment. Two web searches turned up no report from Meta or any other company that quantifies the effect of BOPE in production. In the peer-reviewed literature, hearing aids are the one area with both a described commercial mechanism and user evaluations, and those evaluations involve the manufacturer (inference, from the authors' affiliations; Section 33.4). So the toolchain supports production A/B testing, but the public evidence of such use consists of benchmarks and motivating statements. Dueling bandits that compare search rankers by interleaving results were outside the scope of our searches.

Sources cited in Section 34.3 9
  1. Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
  2. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  3. Meta Platforms, Inc. (2026d) Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1)
  4. Meta Platforms, Inc. (2026l) CHANGELOG (versions 1.2 to 1.3)
  5. Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback
  6. Xu et al. (2026) RankTuner: When Design Tool Parameter Tuning Meets Preference Bayesian Optimization
  7. Sankagiri et al. (2026) Recycling History: Efficient Recommendations from Contextual Dueling Bandits
  8. Huber et al. (2025) Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization
  9. Vyas et al. (2022) Personalizing over-the-counter hearing aids using pairwise comparisons

34.4 The evidence base, domain by domain #

Table 34.1 compares the human evidence across every application domain of Part VII. It counts the studies cited in this part, not the result of a systematic search, and "typical sample" means the number of real people who gave judgments. Its last row collects the domains in which our searches, as of September 2026, turned up no preferential optimization study.

Table 34.1 The human evidence for preferential optimization, by application domain. Counts are of the studies cited in this part of the book.
Domain Studies cited Typical sample (range) Real or simulated people Main limitation
Visual, interface, and generative design (preference feedback) about 24 mostly 10 to 40 (3 to 60, excluding crowdsourcing) mostly real mostly weak comparisons and novices; no test of drift within a session
Rating or performance-based human-in-the-loop optimization about 19 12 to 40 (8 to 200) real not preference feedback; against strong comparisons final quality often no different
Lower-limb exoskeletons (algorithm tuning) 13 (10 with real people), plus 2 self-tuning baselines 2 to 15, median about 5 real, mostly young able-bodied adults only 2 participants with paraplegia; internal validation; almost no comparison with manual tuning
Prostheses 1 PBO (plus 3 related) 2 to 3 amputees real tiny samples; preferences inconsistent across trials
Hearing aids and audio 5 (plus 3 related) 20 to 35 both no improvement in speech clarity; evaluations mostly involve the manufacturer
Visual prostheses 3 17 sighted people; 1 study with unknown sample 1 simulation only, 2 with sighted people no blind users; about 50% agreement between people and the simulated agent
Spinal cord stimulation 3 1 to 5 patients real comparison with physicians qualitative; assumes responses do not change
Thermal comfort and daylight in buildings 8 Purdue samples not obtained; no real people otherwise all simulated since 2024 no trial with real occupants
Driving style and controller calibration 7 (plus 1 non-optimization study) 1 study with 29 drivers; otherwise authors or synthetic decision makers mostly simulated no study on real roads; preferences vary with the scenario
Legged robots and controller tuning about 14 (with method and tool papers) 1 expert or a few lab members few real people an expert's own cost function disagrees with their choices; effort savings not quantified
Active preference-based reward learning 8 about 10 users both mostly linear rewards; Gaussian process versions costly in high dimension
Materials and physical sciences 7 1 to 4 experts more than half simulated experts and non-experts diverge; over-trust in the model
Chemistry and drug discovery 4 35 chemists (MolSkill, not BO); otherwise 1 or simulated mostly simulated only moderate agreement between chemists (kappa 0.32 to 0.40)
Protein design 1 none no human preferences "preference" comes from measured fitness
Industrial platforms and A/B testing 6 no public deployment data synthetic utilities or language-model simulation no public report quantifying production deployment
Domains with no study found 0 not applicable not applicable agriculture; food and flavor with sensory panels; motion sickness and motion-simulator cueing; cochlear implants; deep brain stimulation; functional electrical stimulation in rehabilitation; comfort of prosthetic sockets, seats, and clothing; drones before 2026; driving style on real roads; daylight after 2020; protein design from human experts' preferences; production A/B testing; lasers, welding, and machining (scalar Bayesian optimization only)

Figure 34.1 draws the same table on a common scale.

real peoplereal and simulatedmostly simulatedno human preferences131030100300real people per study (log scale)Visual and generative designRating or performance HITLExoskeletonsProsthesesHearing aids and audioVisual prosthesesSpinal cord stimulationBuildingsno real peopleVehicles and controllersRobot controller tuningPreference reward learningMaterials and physicsChemistry and drugsProtein designno real peopleIndustry and A/B testingno real peopleLower-limb exoskeletons13 studies plus 2 self-tuning baselines · real people · see the wearables and health chapterTypical sample: 2 to 15, median about 5.Main limitation: only 2 participants with paraplegia; internal validation; almost nocomparison with manual tuning.No study found: agriculture; food and flavor with sensory panels; motion sickness; cochlear implants;deep brain stimulation; functional electrical stimulation; comfort of sockets, seats, and clothing;drones before 2026; driving on real roads; daylight after 2020; protein design from human preferences;production A/B testing.
real peoplereal and simulatedmostly simulatedno human preferences131030100300real people per study (log scale)Visual and generative designRating or performance HITLExoskeletonsProsthesesHearing aids and audioVisual prosthesesSpinal cord stimulationBuildingsno real peopleVehicles and controllersRobot controller tuningPreference reward learningMaterials and physicsChemistry and drugsProtein designno real peopleIndustry and A/B testingno real peopleLower-limb exoskeletons13 studies plus 2 self-tuning baselines · realpeople · see the wearables and health chapterTypical sample: 2 to 15, median about 5.Main limitation: only 2 participants withparaplegia; internal validation; almost nocomparison with manual tuning.No study found: agriculture; food and flavor withsensory panels; motion sickness; cochlear implants;deep brain stimulation; functional electricalstimulation; comfort of sockets, seats, andclothing; drones before 2026; driving on realroads; daylight after 2020; protein design fromhuman preferences; production A/B testing.
Figure 34.1 The human evidence by application domain. Each row is a domain of Table 34.1; the thin line spans the number of real people per study, the thick segment the typical range, and a dot marks a single value, median, or study (hollow: not preferential optimization, such as MolSkill). Color says where the judgments came from. Choose a domain to read its study count and main limitation. Values are from the table and approximate where it says so; on the log scale a 35-person study looks close to a 200-person one, so read the numbers in the panel.

The pattern across domains. Outside interactive design, each study with human participants had 1 to 35 people who gave judgments, with a median below 10, and used 12 to 50 comparisons, with one clinical case reaching 564 trials (Section 33.5). Validation agreement was about 65% to 100% and nearly always internal, and comparisons with manual or expert tuning were mostly qualitative. The most rigorously designed studies, the double-blind comparison of Søgaard Jensen et al. (2019) and the repeated blocks of Ingraham et al. (2022) and Arens et al. (2025), are also the ones that report limits: partial benefits, preferences that drift with exposure, a preferred setting that differs from the physiologically best one (inference). Applications also drove the methods: JND-aware acquisition, ordinal and crash labels, multi-objective preference models, and contextual PBO all came out of applications rather than benchmarks (inference).

Three recommendations follow from this evidence (inference), and none of them depends on whether the surrogate is a Gaussian process.

  1. For personalization problems with few parameters (about 4 to 6), where each evaluation is felt with the body and the optimum may be flat, start with self-tuning or a coarse grid search as the baseline, and bring in preferential Bayesian optimization only when that baseline falls short: a thumb-controlled self-tuning reached a 16.6% metabolic reduction in about 11 minutes, with a ±8% tolerance around the preferred timing (Section 33.3).
  2. New studies should include manual tuning, random or coarse grid search, and a linear utility model as comparisons, and measure an objective outcome; comparing only with another preferential optimizer or a slider cannot show whether preferential optimization is needed at all (Section 32.8).
  3. Method papers should test their simulated users on at least a few real people: people agreed with the simulated agent only about 50% of the time in the retinal implant study, and an expert's own cost function disagreed with the expert's choices over most of the range in the robot study (Section 33.5, Section 33.6).

Chapter 46 turns these into practical guidance, and Section 46.9 asks when a simpler method should win.

Sources cited in Section 34.4 3
  1. Søgaard Jensen et al. (2019) Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference
  2. Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
  3. Arens et al. (2025) Preference-based assistance optimization for lifting and lowering with a soft back exosuit

34.5 Common claims, checked #

Secondary accounts of these applications repeat several claims that the primary sources do not support as stated. Table 34.2 checks them.

Table 34.2 Claims about the applications of preferential optimization, checked against the primary sources.
Claim Verdict What the sources show
The exoskeleton gaits of Tucker et al. are among the benchmarks of the qEUBO paper. wrong qEUBO's benchmarks are Ackley, Alpine1, Hartmann, car cab design, Sushi, and an animation task (Astudillo et al., 2023). Simulated exoskeleton personalization appears in the preferential multi-objective paper of Astudillo et al. (Astudillo et al., 2025).
ROIAL characterizes the whole preference landscape. partly right ROIAL learns only within a region of interest that excludes uncomfortable gaits, exploring less than 2% of the action space (Li et al., 2021).
Abdelrahman and Miller (2022) optimize the indoor thermal environment from preference feedback. wrong The paper uses a graph neural network to decide when to collect occupants' thermal preference feedback for personal comfort models; it is not preferential BO and does not optimize the environment (Abdelrahman and Miller, 2022).
Hiranaka et al. (IROS 2023) learn primitive skills from human evaluative feedback. partly right The paper applies reinforcement learning from human feedback over parameterized primitive skills; it does not learn the skills themselves and is not preferential BO (Hiranaka et al., 2023).
BOARS appeared in npj Computational Materials in 2023, CoExBO is a 2023 paper, and BOAP is used for automated science. partly right BOARS was published in 2024 (Biswas et al., 2024); CoExBO appeared at AISTATS 2024 (Adachi et al., 2024); BOAP's expert preferences were simulated from published data (Arun Kumar A V et al., 2024).
Dueling scalarized Thompson sampling is the first provably convergent method for preferential multi-objective optimization, applied to autonomous driving and exoskeletons. partly right Published in TMLR in 2025 (arXiv 2024); both applications are simulated; "first" overlooks the earlier choice-function work on multi-objective optimization of Benavoli et al. (Astudillo et al., 2025; Benavoli et al., 2021b).
PBO works well in experiential domains and where people have expertise, and poorly where preferences must be constructed. an untested synthesis Partial support comes from the split between experts and a non-expert in Mikkola et al. (Mikkola et al., 2020) and from the non-convergence in the field deployment of Ou et al. (Ou et al., 2022). Chapter 45 takes up the question.
Sources cited in Section 34.5 11
  1. Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
  2. Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
  3. Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
  4. Abdelrahman and Miller (2022) Targeting occupant feedback using digital twins: Adaptive spatial-temporal thermal preference sampling to optimize personal comfort models
  5. Hiranaka et al. (2023) Primitive Skill-based Robot Learning from Human Evaluative Feedback
  6. Biswas et al. (2024) A dynamic Bayesian optimized active recommender system for curiosity-driven partially Human-in-the-loop automated experiments
  7. Adachi et al. (2024) Looping in the Human Collaborative and Explainable Bayesian Optimization
  8. Arun Kumar A V et al. (2024) Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties
  9. Benavoli et al. (2021b) Choice functions based multi-objective Bayesian optimisation
  10. Mikkola et al. (2020) Projective Preferential Bayesian Optimization
  11. Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures

34.6 Settled, contested, missing #

Research status Settled, contested, missing

Settled. Every building study of preferential Bayesian optimization since 2024 used simulated occupants. The only large set of expert pairwise judgments, more than 5,000 pairs from 35 chemists, shows moderate agreement between experts (kappa 0.32 to 0.40) and somewhat higher agreement within one expert (about 0.6) (Choung et al., 2023). The software for preference exploration in multi-output experiments exists in BoTorch and Ax. Across domains, studies with real people are small, typically fewer than 10 people outside interactive design, and validated internally.

Contested. Whether expert preferences improve scientific optimization beyond what the measured objective achieves: the positive cases involve one to four experts or simulated ones, and one industrial case found that expert knowledge hurt. How to divide the work between people and algorithms: "human first-computer last" halved costs in one scalar-objective game (Kanarik et al., 2023), but whether the same holds when the objective is a preference is open. Whether experts who are wrong or overconfident can be detected or safely ignored; CoExBO's no-harm guarantee is a partial answer.

Missing. A trial of preference-based HVAC control with real occupants. A driving-style study on real roads. A public, quantified report of preferential optimization in production A/B testing. Any study in agriculture, food and flavor with sensory panels, or protein design from human preferences. A comparison of preferential optimization with measured expert tuning time and quality in controller calibration. Tests of simulated decision makers against real people in method papers.

Sources cited in Section 34.6 2
  1. Choung et al. (2023) Extracting medicinal chemistry intuition via preference machine learning
  2. Kanarik et al. (2023) Human–machine collaboration for improving semiconductor process development

34.7 Exercises #

Exercise 34.1

In each of the two preliminary rounds of MolSkill, the chemists' intra-rater Cohen's kappa on repeated pairs was about 0.6 on average. For a binary choice where each option is chosen half the time, chance agreement is 0.5, and kappa is (po−0.5)/(1−0.5)(p_o - 0.5) / (1 - 0.5), where pop_o is the observed agreement. With that chance level, what fraction of repeated pairs does a chemist with a kappa of 0.6 answer the same way? If you modeled that chemist with the probit likelihood of Chapter 16 and assumed every repeated pair had the same utility difference Δ\Delta (in units of the noise), what Δ\Delta would produce that repeat agreement?

Solution

From 0.6=(po−0.5)/0.50.6 = (p_o - 0.5)/0.5, po=0.8p_o = 0.8: such a chemist gives the same answer to 80% of repeated pairs. Under a probit model each answer favors the better molecule with probability q=Φ(Δ)q = \Phi(\Delta), and two independent answers agree with probability q2+(1−q)2q^2 + (1 - q)^2. Setting q2+(1−q)2=0.8q^2 + (1-q)^2 = 0.8 gives 2q2−2q+0.2=02q^2 - 2q + 0.2 = 0, so q=(1+0.6)/2≈0.887q = (1 + \sqrt{0.6})/2 \approx 0.887, and Δ=Φ−1(0.887)≈1.21\Delta = \Phi^{-1}(0.887) \approx 1.21. In the model's units, a typical repeated pair sits about 1.2 noise standard deviations apart, which is far from a deterministic judge. (The paper's table lists raw repeat agreements of 78.9% to 100%, for example 89.5% beside a kappa of 0.58, which a chance level of 0.5 does not reproduce; with po=0.895p_o = 0.895 the same steps give Δ≈1.59\Delta \approx 1.59, still a noisy judge.) Real pairs differ in Δ\Delta, so this is an average picture, but it shows why a noise-free simulated expert overstates what a real one provides.

Exercise 34.2

POP-BO and qEUBO were compared on a simulated thermal comfort task; qEUBO found a slightly better final setting, while its cumulative regret was almost twice POP-BO's. Explain in one paragraph why a building operator might prefer POP-BO anyway, and describe a deployment in which qEUBO would be the better choice.

Solution

In online HVAC control every query is a room condition that real occupants live in, so each poor query is experienced discomfort. Cumulative regret adds up that discomfort over the whole run, while the final setting matters only after learning ends. An operator who tunes comfort while people work in the building should weigh the cumulative cost, which favors POP-BO. qEUBO would be the better choice when exploration is cheap or happens offline, for example a commissioning phase in an empty test room or with occupants who volunteer for a short calibration session, after which the found setting runs for months. The distinction is the one between simple and cumulative regret in Chapter 13.

Further reading #

  • Choung et al. (2023) is the largest data set of expert pairwise judgments in this chapter and the clearest measurement of how much experts agree.
  • Mikkola et al. (2020) shows on a real physics problem how much the person in the loop matters, with experts and a non-expert reaching different minima.
  • Lin et al. (2022) sets out preference exploration for multi-output experiments, the method behind the industrial tooling.
  • Xu et al. (2024b) includes the thermal comfort simulation and the argument for cumulative regret in online control.
  • Kanarik et al. (2023) is the best evidence on dividing work between people and algorithms, though with a measured objective.

References

  1. Abdelrahman, M., and Miller, C. (2022). Targeting occupant feedback using digital twins: Adaptive spatial-temporal thermal preference sampling to optimize personal comfort models. Building and Environment 218. Cited in §34.5
  2. Adachi, M., Planden, B., Howey, D. A., Osborne, M. A., Orbell, S., Ares, N., Muandet, K., and Chau, S. L. (2024). Looping in the Human Collaborative and Explainable Bayesian Optimization. AISTATS 2024. Cited in §34.2 §34.5
  3. Adachi, M., Chau, S. L., Xu, W., Singh, A., Osborne, M. A., and Muandet, K. (2025). Bayesian Optimization for Building Social-Influence-Free Consensus. arXiv. preprint Cited in §34.1
  4. Arens, P., Quirk, D. A., Pan, W., Yacoby, Y., Doshi-Velez, F., and Walsh, C. J. (2025). Preference-based assistance optimization for lifting and lowering with a soft back exosuit. Science Advances. doi:10.1126/sciadv.adu2099. Cited in §34.4
  5. Arun Kumar A V, Shilton, A., Gupta, S., Rana, S., Greenhill, S., and Venkatesh, S. (2024). Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties. ECML PKDD 2024. Cited in §34.2 §34.5
  6. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §34.3 §34.5
  7. Astudillo, R., Li, K., Tucker, M., Cheng, C. X., Ames, A. D., and Yue, Y. (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Cited in §34.5
  8. Awalgaonkar, N., Bilionis, I., Liu, X., Karava, P., and Tzempelikos, A. (2019). Learning Personalized Thermal Preferences via Bayesian Active Learning with Unimodality Constraints. arXiv. preprint Cited in §34.1
  9. Basu, C., Yang, Q., Hungerman, D., Singhal, M., and Dragan, A. D. (2017). Do You Want Your Autonomous Car To Drive Like You? HRI 2017. Cited in §34.1
  10. Benavoli, A., Azzimonti, D., and Piga, D. (2021b). Choice functions based multi-objective Bayesian optimisation. arXiv. preprint Cited in §34.5
  11. Biswas, A., Liu, Y., Creange, N., Liu, Y.-C., Jesse, S., Yang, J.-C., … Vasudevan, R. K. (2024). A dynamic Bayesian optimized active recommender system for curiosity-driven partially Human-in-the-loop automated experiments. npj Computational Materials. Cited in §34.2 §34.5
  12. Biswas, A., Funakubo, H., and Liu, Y. (2026). Human-AI Collaborative Autonomous Experimentation With Proxy Modeling for Comparative Observation. arXiv. preprint Cited in §34.2
  13. Cercola, M., Lomuscio, M., Piga, D., and Formentin, S. (2026b). Regularized GLISp for sensor-guided human-in-the-loop optimization. IFAC Journal of Systems and Control. doi:10.1016/j.ifacsc.2026.100368. Cited in §34.1
  14. Choung, O.-H., Vianello, R., Segler, M., Stiefl, N., and Jiménez-Luna, J. (2023). Extracting medicinal chemistry intuition via preference machine learning. Nature Communications. doi:10.1038/s41467-023-42242-1. Cited in §34.2 §34.6
  15. Dang, T., Pham, L.-H., Truong, S. T., Glenn, A., Nguyen, W., Pham, E. A., … Luong, T. (2025). Preferential Multi-Objective Bayesian Optimization for Drug Discovery. ICLR 2025 Workshop. workshop paper Cited in §34.2
  16. de Vries, W., van Kampen, J., and Salazar, M. (2024). A Human-optimized Model Predictive Control Scheme and Extremum Seeking Parameter Estimator for Slip Control of Electric Race Cars. arXiv. preprint Cited in §34.1
  17. Deneault, J. R., Kim, W., Kim, J., Gu, Y., Chang, J., Maruyama, B., Myung, J. I., and Pitt, M. A. (2025). Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities. Digital Discovery. Cited in §34.2
  18. Haltia, A., Hyvönen, V., and Kaski, S. (2026). Elicitation-Augmented Bayesian Optimization. arXiv. preprint Cited in §34.2
  19. Hawkins-Hooker, A., Duckworth, P., and Bent, O. (2023). Preferential Bayesian Optimisation for Protein Design with Ranking-Based Fitness Predictors. NeurIPS 2023 MLSB Workshop. workshop paper Cited in §34.2
  20. Hiranaka, A., Hwang, M., Lee, S., Wang, C., Fei-Fei, L., Wu, J., and Zhang, R. (2023). Primitive Skill-based Robot Learning from Human Evaluative Feedback. IROS 2023. Cited in §34.5
  21. Huber, F., Rojas Gonzalez, S., and Astudillo, R. (2025). Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization. Journal of Multi-Criteria Decision Analysis. Cited in §34.3
  22. Ingraham, K. A., Remy, C. D., and Rouse, E. J. (2022). The role of user preference in the customized control of robotic exoskeletons. Science Robotics. Cited in §34.4
  23. Kanarik, K. J., Osowiecki, W. T., Lu, Y., Talukder, D., Roschewsky, N., Park, S. N., … Gottscho, R. A. (2023). Human–machine collaboration for improving semiconductor process development. Nature. Cited in §34.2 §34.6
  24. Kobalczyk, K., Lin, Z. J., Letham, B., Zhao, Z., Balandat, M., and Bakshy, E. (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Cited in §34.3
  25. Kristiadi, A., Strieth-Kalthoff, F., Subramanian, S. G., Fortuin, V., Poupart, P., and Pleiss, G. (2024b). How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization? AABI 2024. workshop paper Cited in §34.2
  26. Li, K., Tucker, M., Bıyık, E., Novoseller, E., Burdick, J. W., Sui, Y., … Ames, A. D. (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Cited in §34.5
  27. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §34.3
  28. Liu, Y., and Kalinin, S. V. (2025). Pareto-Optimal Experimentation: Human-Guided Multi-Objective Bayesian Optimization in Scanning Probe Microscopy. Nano Letters. Cited in §34.2
  29. Meta Platforms, Inc. (2026d). Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1). botorch.org. software Cited in §34.3
  30. Meta Platforms, Inc. (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Cited in §34.3
  31. Mikkola, P., Todorović, M., Järvi, J., Rinke, P., and Kaski, S. (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §34.2 §34.5
  32. Ou, C., Buschek, D., Mayer, S., and Butz, A. (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §34.5
  33. Ran, W., Chen, H., Xia, T., Nishimura, Y., Guo, C., and Yin, Y. (2023). Online Personalized Preference Learning Method Based on In-Formative Query for Lane Centering Control Trajectory. Sensors. doi:10.3390/s23115246. Cited in §34.1
  34. Sankagiri, S., Etesami, J., Fatemi, P., and Grossglauser, M. (2026). Recycling History: Efficient Recommendations from Contextual Dueling Bandits. Algorithmic Learning Theory. Cited in §34.3
  35. Søgaard Jensen, N., Hau, O., Bagger Nielsen, J. B., Bundgaard Nielsen, T., and Vase Legarth, S. (2019). Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference. Trends in Hearing. doi:10.1177/2331216519847413. Cited in §34.4
  36. Sundin, I., Voronov, A., Xiao, H., Papadopoulos, K., Bjerrum, E. J., Heinonen, M., … Engkvist, O. (2022). Human-in-the-loop assisted de novo molecular design. Journal of Cheminformatics. doi:10.1186/s13321-022-00667-8. Cited in §34.2
  37. Theiner, L., Hirt, S., Steinke, A., and Findeisen, R. (2025). Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles. ECC 2025. Cited in §34.1
  38. Theiner, L., Pfefferkorn, M., Zhao, Y., Hirt, S., and Findeisen, R. (2026). Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models. European Control Conference. Cited in §34.1
  39. Vyas, D., Brummet, R., Anwar, Y., Jensen, J., Jorgensen, E., Wu, Y.-H., and Chipara, O. (2022). Personalizing over-the-counter hearing aids using pairwise comparisons. Smart Health. doi:10.1016/j.smhl.2021.100231. Cited in §34.3
  40. Wang, W., Xu, W., and Jones, C. N. (2025c). Human-in-the-loop: Real-time Preference Optimization. arXiv. preprint Cited in §34.1
  41. Wang, W., Shi, J., and Jones, C. N. (2025d). Personalized Building Climate Control with Contextual Preferential Bayesian Optimization. arXiv. preprint Cited in §34.1
  42. Weichert, D., Ernis, G., Worthmann, M., Ryzko, P., and Seifert, L. (2025). When Less is More: A Story of Failing Bayesian Optimization Due to Additional Expert Knowledge. arXiv. preprint Cited in §34.2
  43. Xiong, J., Lee, S., Karava, P., and Tzempelikos, A. (2017). Personalized visual satisfaction profiles from comparative preferences using Bayesian inference. Energy Procedia. Cited in §34.1
  44. Xiong, J., Tzempelikos, A., Bilionis, I., Awalgaonkar, N. M., Lee, S., Konstantzos, I., Sadeghi, S. A., and Karava, P. (2018). Inferring personalized visual satisfaction profiles in daylit offices from comparative preferences using a Bayesian approach. Building and Environment. Cited in §34.1
  45. Xiong, J., Awalgaonkar, N. M., Tzempelikos, A., Bilionis, I., and Karava, P. (2020). Efficient learning of personalized visual preferences in daylit offices: An online elicitation framework. Building and Environment. Cited in §34.1
  46. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §34.1
  47. Xu, P., Zheng, S., Ye, Y., Bai, C., Xu, S., Geng, H., Ho, T.-Y., and Yu, B. (2026). RankTuner: When Design Tool Parameter Tuning Meets Preference Bayesian Optimization. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems. Cited in §34.3
  48. Zhu, M., Bemporad, A., and Piga, D. (2021). Preference-based MPC calibration. ECC 2021. Cited in §34.1
  49. Zhu, M., Piga, D., and Bemporad, A. (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Cited in §34.1