Built Environments, Science, and Industry
Chapter 32 and Chapter 33 followed preferential optimization where people were in the loop by necessity: a design only its user can judge, a device only its wearer can feel. This chapter covers the remaining applications, where the case for asking a person is less obvious and the evidence takes a different shape. In buildings and vehicles, comfort and driving style are personal, but the studies mostly use simulated occupants and drivers. In science and engineering, preferences inject an expert's judgment into an optimization that otherwise has a measured objective, and the experts are usually one to four people. In industry, the software is ready, and the public record consists of benchmarks and motivating statements. The chapter then compares the human evidence across every application domain of this part, and ends by checking claims about these applications that circulate in secondary accounts.
34.1 Buildings and vehicles #
34.1.1 Thermal comfort and daylight #
Comfort differs between people, and occupants answer comparisons about it easily, which makes buildings a natural target. The evidence is almost entirely simulated. A Purdue University group learned individual visual satisfaction with office daylighting from comparative preferences between 2017 and 2020 (Xiong et al., 2017; Xiong et al., 2018; Xiong et al., 2020); we could not obtain the sample sizes or the number of queries. Awalgaonkar et al. (2019), a preprint, treated an occupant's answer ("I would like it warmer", "cooler", or "I am satisfied") as the sign of the derivative of their utility, used a Gaussian process constrained to a single peak, and validated it with synthetic occupants.
POP-BO, an optimistic method of preferential Bayesian optimization (PBO) with a bound on cumulative regret (Xu et al., 2024b), includes a thermal comfort task in simulation, with preferences generated from the predicted mean vote (PMV), a standard formula for a group's average comfort rating, over 2 or 4 variables such as air temperature and air speed. qEUBO (EUBO for queries of options, Section 19.4) found a slightly better final solution, but its cumulative regret was almost twice that of POP-BO, and the authors argue that online control of heating, ventilation, and air conditioning (HVAC) should care about cumulative regret, the total discomfort over the whole run (Chapter 13). Three more recent preprints are simulations too: contextual PBO raised occupant utility by up to 23% over two simulated months on BOPTEST, a building simulation platform (Wang et al., 2025d); a real-time preference-optimizing controller was evaluated on simulated thermal comfort (Wang et al., 2025c); and a consensus Bayesian optimization robust to social influence uses thermal comfort as one application (Adachi et al., 2025).
Every building study of PBO from 2024 to 2026 used simulated occupants. That is a gap, not a negative result: the next useful study is a trial of preference-based HVAC control with real occupants, reporting cumulative discomfort as well as the final setting (inference).
34.1.2 Vehicles and controller calibration #
People misjudge their own driving. In Basu et al. (2017), not an optimization study, users preferred a driving style clearly more defensive than how they drove, preferred the style they believed was their own over their actual one, and preferred different styles in different scenarios, so a system that learns from how people drive and one that learns from what they say they prefer reach different places.
One study with real drivers. Ran et al. (2023) personalized lane-centering trajectories online from preferences, modeling the uncertainty in each answer; in a fixed-base driving simulator, 29 drivers (15 inexperienced, 14 experienced) converged after 11.1 ± 4.6 queries on average, almost all in fewer than 20, and the estimated utility agreed with their own rating scores.
Controller calibration in simulation. A model predictive controller (MPC) chooses each action by optimizing over a short prediction of the future, with weights an engineer usually calibrates by trial and error. Zhu et al. (2021) calibrated MPC controllers with GLISp, the radial basis function method of Section 33.6, on a simulated reactor and a simulated lane-keeping car (50 experiments each); for the car, the authors could not write a scoring function usable for automatic calibration even after much trial and error. In the later studies the decision maker was synthetic or an author. C-GLISp adds feasibility labels for unknown constraints, with a synthetic decision maker in its benchmarks and the first author as calibrator in its driving case (Zhu et al., 2022); Theiner et al. (2025) guided PBO with a virtual decision maker built from real driving data, again with a synthetic main decision maker, and converged faster than standard PBO, which needed about 70 iterations, with a multi-fidelity follow-up (Theiner et al., 2026); and two preprints tuned a vehicle suspension with a synthetic user (Cercola et al., 2026b) and the slip control of an electric race car, with simulation results only (de Vries et al., 2024). The in-car interface studies optimized from ratings are in Section 32.1.4.
Sources cited in Section 34.1 16
- Xiong et al. (2017) Personalized visual satisfaction profiles from comparative preferences using Bayesian inference
- Xiong et al. (2018) Inferring personalized visual satisfaction profiles in daylit offices from comparative preferences using a Bayesian approach
- Xiong et al. (2020) Efficient learning of personalized visual preferences in daylit offices: An online elicitation framework
- Awalgaonkar et al. (2019) Learning Personalized Thermal Preferences via Bayesian Active Learning with Unimodality Constraints
- Xu et al. (2024b) Principled Preferential Bayesian Optimization
- Wang et al. (2025d) Personalized Building Climate Control with Contextual Preferential Bayesian Optimization
- Wang et al. (2025c) Human-in-the-loop: Real-time Preference Optimization
- Adachi et al. (2025) Bayesian Optimization for Building Social-Influence-Free Consensus
- Basu et al. (2017) Do You Want Your Autonomous Car To Drive Like You?
- Ran et al. (2023) Online Personalized Preference Learning Method Based on In-Formative Query for Lane Centering Control Trajectory
- Zhu et al. (2021) Preference-based MPC calibration
- Zhu et al. (2022) C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration
- Theiner et al. (2025) Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles
- Theiner et al. (2026) Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models
- Cercola et al. (2026b) Regularized GLISp for sensor-guided human-in-the-loop optimization
- de Vries et al. (2024) A Human-optimized Model Predictive Control Scheme and Extremum Seeking Parameter Estimator for Slip Control of Electric Race Cars
34.2 Science and engineering with expert preferences #
In the sciences, preferences mostly bring an expert's judgment into a Bayesian optimization whose objective would otherwise be a measured scalar, and in each study the real experts number one to four. Section 15.3 covers Bayesian optimization in experimental science with measured objectives, and Chapter 23 works through a chemical reaction end to end.
34.2.1 Materials and physical sciences #
Expertise visibly changes the outcome. The projective preferential Bayesian optimization of Mikkola et al. (2020) asks the user to pick the best point along a one-dimensional projection with a slider (Section 20.3). It was used to find the position and orientation of a camphor molecule on a copper surface, Cu(111), with each candidate's energy checked by density functional theory, a quantum-mechanical calculation of the energy of a molecular configuration. Every user answered 24 queries. The structures preferred by two materials-science experts relaxed to energy minima of about -1.007 to -1.030 eV, while those of a non-expert relaxed to shallower local minima of about -0.762 to -0.771 eV. The authors describe this as a clear split by expertise, and the method could also tell a person's choices from those of a random robot.
Votes on spectra. BOARS turns an operator's up or down votes on measured spectra into an objective by Thurstone-Mosteller scaling (Section 16.3), with the human leading early and the algorithm later, demonstrated on a live atomic force microscope (Biswas et al., 2024). A 2026 preprint, px-BO, fits a Bradley-Terry model (Section 16.4) to the votes and then lets a surrogate vote in the human's place, with periodic human checks (Biswas et al., 2026).
Explanations, and the risk of over-trust. In CoExBO (Section 32.7), an expert chooses between two candidates shown with explanations (Adachi et al., 2024). On battery electrolyte design with 4 participants, the explanations improved the experts' accuracy; the authors note that the gains assume the expert's comparative knowledge is accurate, and that experts tended to expect the surrogate to understand the problem like an oracle, a form of over-trust.
Simulated experts. BOAP models an expert's pairwise preferences over "abstract properties" that cannot be measured, for lithium-ion electrode manufacturing, but the preferences were simulated from published data sets (Arun Kumar A V et al., 2024). In the multi-objective optimization of scanning probe microscopy by Liu and Kalinin (2025), the human steers by objective weights and reference points, not comparisons. Deneault et al. (2025) used PBO to optimize subjective qualities of 3-D prints, judged by a person rather than measured by sensors.
34.2.2 Chemistry and drug discovery #
Chemistry holds the only large data set of expert preferences in this chapter. MolSkill collected more than 5,000 pairwise comparisons of molecules from 35 chemists at Novartis (in wet-lab, computational, and analytical roles) over several months, choosing each round by active learning (Choung et al., 2023). Agreement was moderate. In preliminary rounds, the inter-rater agreement on 200 pairs, measured by Fleiss' kappa, a chance-corrected agreement among several raters where 0 is chance and 1 is perfect, was 0.40 and 0.32; the intra-rater agreement on 20 repeated pairs, measured by Cohen's kappa, was 0.60 and 0.59, and for some chemists as low as 0.16 to 0.27. The area under the ROC curve for classifying pairs, the probability that the model ranks a random preferred molecule above a random non-preferred one, rose from about 0.6 with 1,000 pairs to above 0.74 with 5,000, without reaching a plateau. The learned score captured drug-likeness better than QED, the quantitative estimate of drug-likeness chemists commonly use. MolSkill is a neural network that learns to rank, not Bayesian optimization.
The other chemistry studies use simulated experts, a single expert, or none. Sundin et al. (2022) learned a chemist's scoring function actively, improving significantly in fewer than 200 queries in two cases with simulated ground truth, with one medicinal chemist in a demonstration who judged one molecule at a time, not pairs. CheapVS, a workshop paper, guided screening of 100,000 molecules with chemists' pairwise preferences over trade-offs among properties: screening 6% of the library recovered 16 of the 37 known drugs for one target (EGFR) and 37 of the 58 for another (DRD2), but with a synthetic utility function, preliminary human data (expert rankings for EGFR about 80% accurate), and a number of experts we could not find (Dang et al., 2025). Kristiadi et al. (2024b), a workshop paper, simulated experts' pairwise labels with a scoring function; the "preferential Bayesian optimization" for protein design of Hawkins-Hooker et al. (2023), also a workshop paper, takes its comparisons from measured fitness; and Haltia et al. (2026), a preprint, choose between an expensive evaluation and a query to an expert by their costs.
34.2.3 Dividing the work between people and algorithms #
Kanarik et al. (2023) compared people and Bayesian optimization in a virtual process game for designing a semiconductor plasma etching process. Engineers did well early, the algorithm was far more cost-efficient close to tight tolerances, and a "human first-computer last" strategy halved the cost of reaching the target compared with engineers alone. The objective was a measured scalar, so how far this carries over to preference-driven design is uncertain. And Weichert et al. (2025), a preprint, document an industrial case in which adding expert knowledge made the optimization fail.
What it means. Agreement of about 0.3 to 0.4 between chemists and about 0.6 within one says that experts' pairwise intuition can be learned but is only moderately shared, which favors models per expert or mixture models over pooled data (inference; Section 27.2). Mikkola et al.'s split between experts and a non-expert, and CoExBO's assumption that expert knowledge is accurate, point to an untested failure mode: the expert who is wrong or overconfident. CoExBO's no-harm guarantee is the main design response so far (inference), and a method that relies on expert preferences could be tested with at least one deliberately misinformed expert (inference).
Sources cited in Section 34.2 15
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Biswas et al. (2024) A dynamic Bayesian optimized active recommender system for curiosity-driven partially Human-in-the-loop automated experiments
- Biswas et al. (2026) Human-AI Collaborative Autonomous Experimentation With Proxy Modeling for Comparative Observation
- Adachi et al. (2024) Looping in the Human Collaborative and Explainable Bayesian Optimization
- Arun Kumar A V et al. (2024) Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties
- Liu and Kalinin (2025) Pareto-Optimal Experimentation: Human-Guided Multi-Objective Bayesian Optimization in Scanning Probe Microscopy
- Deneault et al. (2025) Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities
- Choung et al. (2023) Extracting medicinal chemistry intuition via preference machine learning
- Sundin et al. (2022) Human-in-the-loop assisted de novo molecular design
- Dang et al. (2025) Preferential Multi-Objective Bayesian Optimization for Drug Discovery
- Kristiadi et al. (2024b) How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization?
- Hawkins-Hooker et al. (2023) Preferential Bayesian Optimisation for Protein Design with Ranking-Based Fitness Predictors
- Haltia et al. (2026) Elicitation-Augmented Bayesian Optimization
- Kanarik et al. (2023) Human–machine collaboration for improving semiconductor process development
- Weichert et al. (2025) When Less is More: A Story of Failing Bayesian Optimization Due to Additional Expert Knowledge
34.3 Industry #
The main industrial line: preference exploration at Meta. BOPE (preference exploration for Bayesian optimization with multiple outcomes) alternates between running experiments and asking a decision maker to compare predicted outcome vectors (Lin et al., 2022). The paper is motivated by A/B testing workflows, but its experiments use a vehicle safety problem (5 inputs, 3 outputs), a car cab design problem (7 inputs, 9 outputs), and the DTLZ2 and OSY test problems, with a noisy synthetic utility standing in for a human decision maker. qEUBO extended the expected utility of the best option to queries of options (Astudillo et al., 2023). Both are in BoTorch (Meta Platforms, Inc., 2026d), and between January and June 2026 Ax versions 1.2 to 1.3 added preference optimization configuration, BOPE utility tracking, qEUBO, and pairwise Gaussian processes (Meta Platforms, Inc., 2026l) (Section 31.1.2). LILO, which turns natural-language feedback into preference signals with a language model, was evaluated with a language model standing in for the decision maker (Kobalczyk et al., 2026) (Chapter 35). Other industry-facing applications include banner ads (Section 32.1.1), RankTuner for electronic design automation tools (Xu et al., 2026), contextual dueling bandits for recommendation (Sankagiri et al., 2026), multi-criteria decision support (Huber et al., 2025), and hearing-aid presets (Vyas et al., 2022).
Public evidence of deployment. Two web searches turned up no report from Meta or any other company that quantifies the effect of BOPE in production. In the peer-reviewed literature, hearing aids are the one area with both a described commercial mechanism and user evaluations, and those evaluations involve the manufacturer (inference, from the authors' affiliations; Section 33.4). So the toolchain supports production A/B testing, but the public evidence of such use consists of benchmarks and motivating statements. Dueling bandits that compare search rankers by interleaving results were outside the scope of our searches.
Sources cited in Section 34.3 9
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Meta Platforms, Inc. (2026d) Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1)
- Meta Platforms, Inc. (2026l) CHANGELOG (versions 1.2 to 1.3)
- Kobalczyk et al. (2026) LILO: Bayesian Optimization with Natural Language Feedback
- Xu et al. (2026) RankTuner: When Design Tool Parameter Tuning Meets Preference Bayesian Optimization
- Sankagiri et al. (2026) Recycling History: Efficient Recommendations from Contextual Dueling Bandits
- Huber et al. (2025) Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization
- Vyas et al. (2022) Personalizing over-the-counter hearing aids using pairwise comparisons
34.4 The evidence base, domain by domain #
Table 34.1 compares the human evidence across every application domain of Part VII. It counts the studies cited in this part, not the result of a systematic search, and "typical sample" means the number of real people who gave judgments. Its last row collects the domains in which our searches, as of September 2026, turned up no preferential optimization study.
| Domain | Studies cited | Typical sample (range) | Real or simulated people | Main limitation |
|---|---|---|---|---|
| Visual, interface, and generative design (preference feedback) | about 24 | mostly 10 to 40 (3 to 60, excluding crowdsourcing) | mostly real | mostly weak comparisons and novices; no test of drift within a session |
| Rating or performance-based human-in-the-loop optimization | about 19 | 12 to 40 (8 to 200) | real | not preference feedback; against strong comparisons final quality often no different |
| Lower-limb exoskeletons (algorithm tuning) | 13 (10 with real people), plus 2 self-tuning baselines | 2 to 15, median about 5 | real, mostly young able-bodied adults | only 2 participants with paraplegia; internal validation; almost no comparison with manual tuning |
| Prostheses | 1 PBO (plus 3 related) | 2 to 3 amputees | real | tiny samples; preferences inconsistent across trials |
| Hearing aids and audio | 5 (plus 3 related) | 20 to 35 | both | no improvement in speech clarity; evaluations mostly involve the manufacturer |
| Visual prostheses | 3 | 17 sighted people; 1 study with unknown sample | 1 simulation only, 2 with sighted people | no blind users; about 50% agreement between people and the simulated agent |
| Spinal cord stimulation | 3 | 1 to 5 patients | real | comparison with physicians qualitative; assumes responses do not change |
| Thermal comfort and daylight in buildings | 8 | Purdue samples not obtained; no real people otherwise | all simulated since 2024 | no trial with real occupants |
| Driving style and controller calibration | 7 (plus 1 non-optimization study) | 1 study with 29 drivers; otherwise authors or synthetic decision makers | mostly simulated | no study on real roads; preferences vary with the scenario |
| Legged robots and controller tuning | about 14 (with method and tool papers) | 1 expert or a few lab members | few real people | an expert's own cost function disagrees with their choices; effort savings not quantified |
| Active preference-based reward learning | 8 | about 10 users | both | mostly linear rewards; Gaussian process versions costly in high dimension |
| Materials and physical sciences | 7 | 1 to 4 experts | more than half simulated | experts and non-experts diverge; over-trust in the model |
| Chemistry and drug discovery | 4 | 35 chemists (MolSkill, not BO); otherwise 1 or simulated | mostly simulated | only moderate agreement between chemists (kappa 0.32 to 0.40) |
| Protein design | 1 | none | no human preferences | "preference" comes from measured fitness |
| Industrial platforms and A/B testing | 6 | no public deployment data | synthetic utilities or language-model simulation | no public report quantifying production deployment |
| Domains with no study found | 0 | not applicable | not applicable | agriculture; food and flavor with sensory panels; motion sickness and motion-simulator cueing; cochlear implants; deep brain stimulation; functional electrical stimulation in rehabilitation; comfort of prosthetic sockets, seats, and clothing; drones before 2026; driving style on real roads; daylight after 2020; protein design from human experts' preferences; production A/B testing; lasers, welding, and machining (scalar Bayesian optimization only) |
Figure 34.1 draws the same table on a common scale.
The pattern across domains. Outside interactive design, each study with human participants had 1 to 35 people who gave judgments, with a median below 10, and used 12 to 50 comparisons, with one clinical case reaching 564 trials (Section 33.5). Validation agreement was about 65% to 100% and nearly always internal, and comparisons with manual or expert tuning were mostly qualitative. The most rigorously designed studies, the double-blind comparison of Søgaard Jensen et al. (2019) and the repeated blocks of Ingraham et al. (2022) and Arens et al. (2025), are also the ones that report limits: partial benefits, preferences that drift with exposure, a preferred setting that differs from the physiologically best one (inference). Applications also drove the methods: JND-aware acquisition, ordinal and crash labels, multi-objective preference models, and contextual PBO all came out of applications rather than benchmarks (inference).
Three recommendations follow from this evidence (inference), and none of them depends on whether the surrogate is a Gaussian process.
- For personalization problems with few parameters (about 4 to 6), where each evaluation is felt with the body and the optimum may be flat, start with self-tuning or a coarse grid search as the baseline, and bring in preferential Bayesian optimization only when that baseline falls short: a thumb-controlled self-tuning reached a 16.6% metabolic reduction in about 11 minutes, with a ±8% tolerance around the preferred timing (Section 33.3).
- New studies should include manual tuning, random or coarse grid search, and a linear utility model as comparisons, and measure an objective outcome; comparing only with another preferential optimizer or a slider cannot show whether preferential optimization is needed at all (Section 32.8).
- Method papers should test their simulated users on at least a few real people: people agreed with the simulated agent only about 50% of the time in the retinal implant study, and an expert's own cost function disagreed with the expert's choices over most of the range in the robot study (Section 33.5, Section 33.6).
Chapter 46 turns these into practical guidance, and Section 46.9 asks when a simpler method should win.
Sources cited in Section 34.4 3
- Søgaard Jensen et al. (2019) Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference
- Ingraham et al. (2022) The role of user preference in the customized control of robotic exoskeletons
- Arens et al. (2025) Preference-based assistance optimization for lifting and lowering with a soft back exosuit
34.5 Common claims, checked #
Secondary accounts of these applications repeat several claims that the primary sources do not support as stated. Table 34.2 checks them.
| Claim | Verdict | What the sources show |
|---|---|---|
| The exoskeleton gaits of Tucker et al. are among the benchmarks of the qEUBO paper. | wrong | qEUBO's benchmarks are Ackley, Alpine1, Hartmann, car cab design, Sushi, and an animation task (Astudillo et al., 2023). Simulated exoskeleton personalization appears in the preferential multi-objective paper of Astudillo et al. (Astudillo et al., 2025). |
| ROIAL characterizes the whole preference landscape. | partly right | ROIAL learns only within a region of interest that excludes uncomfortable gaits, exploring less than 2% of the action space (Li et al., 2021). |
| Abdelrahman and Miller (2022) optimize the indoor thermal environment from preference feedback. | wrong | The paper uses a graph neural network to decide when to collect occupants' thermal preference feedback for personal comfort models; it is not preferential BO and does not optimize the environment (Abdelrahman and Miller, 2022). |
| Hiranaka et al. (IROS 2023) learn primitive skills from human evaluative feedback. | partly right | The paper applies reinforcement learning from human feedback over parameterized primitive skills; it does not learn the skills themselves and is not preferential BO (Hiranaka et al., 2023). |
| BOARS appeared in npj Computational Materials in 2023, CoExBO is a 2023 paper, and BOAP is used for automated science. | partly right | BOARS was published in 2024 (Biswas et al., 2024); CoExBO appeared at AISTATS 2024 (Adachi et al., 2024); BOAP's expert preferences were simulated from published data (Arun Kumar A V et al., 2024). |
| Dueling scalarized Thompson sampling is the first provably convergent method for preferential multi-objective optimization, applied to autonomous driving and exoskeletons. | partly right | Published in TMLR in 2025 (arXiv 2024); both applications are simulated; "first" overlooks the earlier choice-function work on multi-objective optimization of Benavoli et al. (Astudillo et al., 2025; Benavoli et al., 2021b). |
| PBO works well in experiential domains and where people have expertise, and poorly where preferences must be constructed. | an untested synthesis | Partial support comes from the split between experts and a non-expert in Mikkola et al. (Mikkola et al., 2020) and from the non-convergence in the field deployment of Ou et al. (Ou et al., 2022). Chapter 45 takes up the question. |
Sources cited in Section 34.5 11
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
- Astudillo et al. (2025) Preferential Multi-Objective Bayesian Optimization
- Li et al. (2021) ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
- Abdelrahman and Miller (2022) Targeting occupant feedback using digital twins: Adaptive spatial-temporal thermal preference sampling to optimize personal comfort models
- Hiranaka et al. (2023) Primitive Skill-based Robot Learning from Human Evaluative Feedback
- Biswas et al. (2024) A dynamic Bayesian optimized active recommender system for curiosity-driven partially Human-in-the-loop automated experiments
- Adachi et al. (2024) Looping in the Human Collaborative and Explainable Bayesian Optimization
- Arun Kumar A V et al. (2024) Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties
- Benavoli et al. (2021b) Choice functions based multi-objective Bayesian optimisation
- Mikkola et al. (2020) Projective Preferential Bayesian Optimization
- Ou et al. (2022) The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures
34.6 Settled, contested, missing #
Settled. Every building study of preferential Bayesian optimization since 2024 used simulated occupants. The only large set of expert pairwise judgments, more than 5,000 pairs from 35 chemists, shows moderate agreement between experts (kappa 0.32 to 0.40) and somewhat higher agreement within one expert (about 0.6) (Choung et al., 2023). The software for preference exploration in multi-output experiments exists in BoTorch and Ax. Across domains, studies with real people are small, typically fewer than 10 people outside interactive design, and validated internally.
Contested. Whether expert preferences improve scientific optimization beyond what the measured objective achieves: the positive cases involve one to four experts or simulated ones, and one industrial case found that expert knowledge hurt. How to divide the work between people and algorithms: "human first-computer last" halved costs in one scalar-objective game (Kanarik et al., 2023), but whether the same holds when the objective is a preference is open. Whether experts who are wrong or overconfident can be detected or safely ignored; CoExBO's no-harm guarantee is a partial answer.
Missing. A trial of preference-based HVAC control with real occupants. A driving-style study on real roads. A public, quantified report of preferential optimization in production A/B testing. Any study in agriculture, food and flavor with sensory panels, or protein design from human preferences. A comparison of preferential optimization with measured expert tuning time and quality in controller calibration. Tests of simulated decision makers against real people in method papers.
Sources cited in Section 34.6 2
- Choung et al. (2023) Extracting medicinal chemistry intuition via preference machine learning
- Kanarik et al. (2023) Human–machine collaboration for improving semiconductor process development
34.7 Exercises #
In each of the two preliminary rounds of MolSkill, the chemists' intra-rater Cohen's kappa on repeated pairs was about 0.6 on average. For a binary choice where each option is chosen half the time, chance agreement is 0.5, and kappa is , where is the observed agreement. With that chance level, what fraction of repeated pairs does a chemist with a kappa of 0.6 answer the same way? If you modeled that chemist with the probit likelihood of Chapter 16 and assumed every repeated pair had the same utility difference (in units of the noise), what would produce that repeat agreement?
Solution
From , : such a chemist gives the same answer to 80% of repeated pairs. Under a probit model each answer favors the better molecule with probability , and two independent answers agree with probability . Setting gives , so , and . In the model's units, a typical repeated pair sits about 1.2 noise standard deviations apart, which is far from a deterministic judge. (The paper's table lists raw repeat agreements of 78.9% to 100%, for example 89.5% beside a kappa of 0.58, which a chance level of 0.5 does not reproduce; with the same steps give , still a noisy judge.) Real pairs differ in , so this is an average picture, but it shows why a noise-free simulated expert overstates what a real one provides.
POP-BO and qEUBO were compared on a simulated thermal comfort task; qEUBO found a slightly better final setting, while its cumulative regret was almost twice POP-BO's. Explain in one paragraph why a building operator might prefer POP-BO anyway, and describe a deployment in which qEUBO would be the better choice.
Solution
In online HVAC control every query is a room condition that real occupants live in, so each poor query is experienced discomfort. Cumulative regret adds up that discomfort over the whole run, while the final setting matters only after learning ends. An operator who tunes comfort while people work in the building should weigh the cumulative cost, which favors POP-BO. qEUBO would be the better choice when exploration is cheap or happens offline, for example a commissioning phase in an empty test room or with occupants who volunteer for a short calibration session, after which the found setting runs for months. The distinction is the one between simple and cumulative regret in Chapter 13.
Further reading #
- Choung et al. (2023) is the largest data set of expert pairwise judgments in this chapter and the clearest measurement of how much experts agree.
- Mikkola et al. (2020) shows on a real physics problem how much the person in the loop matters, with experts and a non-expert reaching different minima.
- Lin et al. (2022) sets out preference exploration for multi-output experiments, the method behind the industrial tooling.
- Xu et al. (2024b) includes the thermal comfort simulation and the argument for cumulative regret in online control.
- Kanarik et al. (2023) is the best evidence on dividing work between people and algorithms, though with a measured objective.
References
- (2022). Targeting occupant feedback using digital twins: Adaptive spatial-temporal thermal preference sampling to optimize personal comfort models. Building and Environment 218. Cited in §34.5
- (2024). Looping in the Human Collaborative and Explainable Bayesian Optimization. AISTATS 2024. Cited in §34.2 §34.5
- (2025). Bayesian Optimization for Building Social-Influence-Free Consensus. arXiv. preprint Cited in §34.1
- (2025). Preference-based assistance optimization for lifting and lowering with a soft back exosuit. Science Advances. doi:10.1126/sciadv.adu2099. Cited in §34.4
- (2024). Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties. ECML PKDD 2024. Cited in §34.2 §34.5
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §34.3 §34.5
- (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Cited in §34.5
- (2019). Learning Personalized Thermal Preferences via Bayesian Active Learning with Unimodality Constraints. arXiv. preprint Cited in §34.1
- (2017). Do You Want Your Autonomous Car To Drive Like You? HRI 2017. Cited in §34.1
- (2024). A dynamic Bayesian optimized active recommender system for curiosity-driven partially Human-in-the-loop automated experiments. npj Computational Materials. Cited in §34.2 §34.5
- (2026). Human-AI Collaborative Autonomous Experimentation With Proxy Modeling for Comparative Observation. arXiv. preprint Cited in §34.2
- (2026b). Regularized GLISp for sensor-guided human-in-the-loop optimization. IFAC Journal of Systems and Control. doi:10.1016/j.ifacsc.2026.100368. Cited in §34.1
- (2023). Extracting medicinal chemistry intuition via preference machine learning. Nature Communications. doi:10.1038/s41467-023-42242-1. Cited in §34.2 §34.6
- (2025). Preferential Multi-Objective Bayesian Optimization for Drug Discovery. ICLR 2025 Workshop. workshop paper Cited in §34.2
- (2024). A Human-optimized Model Predictive Control Scheme and Extremum Seeking Parameter Estimator for Slip Control of Electric Race Cars. arXiv. preprint Cited in §34.1
- (2025). Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities. Digital Discovery. Cited in §34.2
- (2026). Elicitation-Augmented Bayesian Optimization. arXiv. preprint Cited in §34.2
- (2023). Preferential Bayesian Optimisation for Protein Design with Ranking-Based Fitness Predictors. NeurIPS 2023 MLSB Workshop. workshop paper Cited in §34.2
- (2023). Primitive Skill-based Robot Learning from Human Evaluative Feedback. IROS 2023. Cited in §34.5
- (2025). Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization. Journal of Multi-Criteria Decision Analysis. Cited in §34.3
- (2022). The role of user preference in the customized control of robotic exoskeletons. Science Robotics. Cited in §34.4
- (2023). Human–machine collaboration for improving semiconductor process development. Nature. Cited in §34.2 §34.6
- (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Cited in §34.3
- (2024b). How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization? AABI 2024. workshop paper Cited in §34.2
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Cited in §34.5
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §34.3
- (2025). Pareto-Optimal Experimentation: Human-Guided Multi-Objective Bayesian Optimization in Scanning Probe Microscopy. Nano Letters. Cited in §34.2
- (2026d). Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1). botorch.org. software Cited in §34.3
- (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Cited in §34.3
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §34.2 §34.5
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Cited in §34.5
- (2023). Online Personalized Preference Learning Method Based on In-Formative Query for Lane Centering Control Trajectory. Sensors. doi:10.3390/s23115246. Cited in §34.1
- (2026). Recycling History: Efficient Recommendations from Contextual Dueling Bandits. Algorithmic Learning Theory. Cited in §34.3
- (2019). Perceptual Effects of Adjusting Hearing-Aid Gain by Means of a Machine-Learning Approach Based on Individual User Preference. Trends in Hearing. doi:10.1177/2331216519847413. Cited in §34.4
- (2022). Human-in-the-loop assisted de novo molecular design. Journal of Cheminformatics. doi:10.1186/s13321-022-00667-8. Cited in §34.2
- (2025). Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles. ECC 2025. Cited in §34.1
- (2026). Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models. European Control Conference. Cited in §34.1
- (2022). Personalizing over-the-counter hearing aids using pairwise comparisons. Smart Health. doi:10.1016/j.smhl.2021.100231. Cited in §34.3
- (2025c). Human-in-the-loop: Real-time Preference Optimization. arXiv. preprint Cited in §34.1
- (2025d). Personalized Building Climate Control with Contextual Preferential Bayesian Optimization. arXiv. preprint Cited in §34.1
- (2025). When Less is More: A Story of Failing Bayesian Optimization Due to Additional Expert Knowledge. arXiv. preprint Cited in §34.2
- (2017). Personalized visual satisfaction profiles from comparative preferences using Bayesian inference. Energy Procedia. Cited in §34.1
- (2018). Inferring personalized visual satisfaction profiles in daylit offices from comparative preferences using a Bayesian approach. Building and Environment. Cited in §34.1
- (2020). Efficient learning of personalized visual preferences in daylit offices: An online elicitation framework. Building and Environment. Cited in §34.1
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Cited in §34.1
- (2026). RankTuner: When Design Tool Parameter Tuning Meets Preference Bayesian Optimization. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems. Cited in §34.3
- (2021). Preference-based MPC calibration. ECC 2021. Cited in §34.1
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Cited in §34.1