Where Bayesian Optimization Works
The last four chapters built a method. This one asks where it earns its keep. Bayesian optimization is not the best optimizer for most problems: if you can evaluate the objective a million times, or compute its gradient, simpler tools do better. It pays off under a particular combination of circumstances, and the fields that adopted it are the ones where that combination is common.
The chapter is a tour, not a set of worked examples. Four problems are worked end to end in Part V: tuning a classifier (Chapter 22), optimizing a chemical reaction (Chapter 23), tuning an exoskeleton with a person in the loop (Chapter 24), and enhancing a photo by comparison (Chapter 25). Here each field gets a few paragraphs: what an evaluation is, what makes the field a good or awkward fit, and what the published results show. The tour ends with people, where the evaluation is a human judgment and the next part of the book begins. The last section places Bayesian optimization among the methods it is most often confused with.
15.1 When it pays off #
Four conditions make Bayesian optimization worth its overhead.
- Evaluations are expensive. Each one costs minutes, hours, money, material, or a person's patience, so that the budget is tens or hundreds of evaluations and not millions. The cost of fitting a model and maximizing an acquisition function, a few seconds, is then negligible.
- There is no gradient and no formula. The objective is a black box: a training run, an experiment, a simulation, a judgment.
- There are few inputs, or few that matter. Up to about twenty with standard methods, more with the techniques of Section 14.6.
- Noise is tolerable, and the function has some regularity. Nearby inputs give similar outputs, so a model can generalize from a few evaluations.
When one of these fails, a neighbor usually takes over (Section 15.8). Table 15.1 lists applications where they hold, with what one evaluation is and what the cited study reported.
| Problem | One evaluation is | What the study reported | Source |
|---|---|---|---|
| Tuning the Go program AlphaGo | a set of self-play games | win rate in self-play up from 50% to 66.5% before a match | Chen et al. (2018) (preprint) |
| Conditions of a chemical reaction | one reaction, in batches of five | better average efficiency and consistency than 50 expert chemists and engineers | Shields et al. (2021) |
| Photocatalyst mixtures | one robot-run experiment | 688 experiments over eight days in a ten-variable space; mixtures six times more active | Burger et al. (2020) |
| Fast-charging protocols for batteries | a cycling test, shortened by early prediction | high-cycle-life protocols among 224 candidates in 16 days | Attia et al. (2020) |
| Gait of a quadruped robot | a walking trial | dramatically fewer gait evaluations than local gradient methods | Lizotte et al. (2007) |
| Controller of a quadrotor | a flight | safe, automatic tuning without human intervention | Berkenkamp et al. (2016) |
| Magnets of a free-electron laser | a beam measurement | significantly better than the facility's existing optimizers | Duris et al. (2020) |
| A ranking system at Facebook | a randomized online experiment | demonstrated on live experiments; outperformed existing methods on noisy, constrained synthetic problems | Letham et al. (2019) |
| Timing of a soft exosuit | minutes of walking with measured energy cost | optimum found in 21.4 ± 1.0 minutes; metabolic cost down 17.4 ± 3.2% | Ding et al. (2018) |
| A cookie recipe | a batch baked and rated by tasters | "The cookies improved significantly over time" | Golovin et al. (2017) |
Sources cited in Section 15.1 10
- Chen et al. (2018) Bayesian Optimization in AlphaGo
- Shields et al. (2021) Bayesian reaction optimization as a tool for chemical synthesis
- Burger et al. (2020) A Mobile Robotic Chemist
- Attia et al. (2020) Closed-Loop Optimization of Fast-Charging Protocols for Batteries with Machine Learning
- Lizotte et al. (2007) Automatic Gait Optimization with Gaussian Process Regression
- Berkenkamp et al. (2016) Safe Controller Optimization for Quadrotors with Gaussian Processes
- Duris et al. (2020) Bayesian Optimization of a Free-Electron Laser
- Letham et al. (2019) Constrained Bayesian Optimization with Noisy Experiments
- Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
- Golovin et al. (2017) Google Vizier: A Service for Black-Box Optimization
15.2 Hyperparameter tuning and AutoML #
Machine learning models have settings that training does not choose: learning rates, regularization strengths, the depth of a tree, the number of layers. Each candidate setting must be evaluated by training a model and measuring its error on held-out data, which takes minutes to days. This is the application that made Bayesian optimization widely known. Snoek et al. (2012) showed that with a suitable kernel and a careful treatment of the model's own hyperparameters, it could reach or surpass human experts in tuning models such as convolutional networks. Before that, Bergstra and Bengio (2012) had shown that random search beats grid search when only a few hyperparameters matter, and Bergstra et al. (2011) had proposed sequential model-based alternatives; the question since then has been how much a model of the objective adds to random sampling.
The evidence that it adds something is broad. In the black-box optimization challenge held at NeurIPS 2020, which tuned standard machine learning models on real data sets with held-out objective functions, 61 of 65 teams beat random search, and the best submissions needed over 100 times fewer evaluations to match it (Turner et al., 2021). During the development of the Go program AlphaGo, a 2018 preprint by its developers reports, its hyperparameters were tuned with Bayesian optimization many times; one such tuning before the match with Lee Sedol raised the win rate in self-play games from 50% to 66.5% (Chen et al., 2018). Google's internal service Vizier, described by its developers as "the de facto parameter tuning engine at Google", offers Bayesian optimization among its algorithms (Golovin et al., 2017).
Tuning hyperparameters is one step of a larger automation. Automated machine learning (AutoML) treats the choice of algorithm as one more hyperparameter: Auto-WEKA searched jointly over 27 base classifiers, their settings, and feature selection methods, a space with many categorical and conditional choices, with Bayesian optimization methods built for such spaces (Thornton et al., 2013). Conditional spaces (the number of trees matters only if the algorithm is a forest) are a reason such systems often prefer tree-based surrogates to Gaussian processes (Section 14.1.2).
Two features of this field shaped the methods of Chapter 14. Evaluations vary enormously in cost, and cheaper approximations are available, so multi-fidelity methods matter (Section 14.7.2). And evaluations run in parallel on clusters, so batches matter (Section 14.3).
A caution belongs here too. On a small problem, the gain over random search can be modest. Chapter 22 measures the full landscape of a two-hyperparameter problem and finds that Bayesian optimization and random search reach nearly the same median error; what the model buys is reliability, and the gap in the median opens only with seven hyperparameters (Section 22.4).
Sources cited in Section 15.2 7
- Snoek et al. (2012) Practical Bayesian Optimization of Machine Learning Algorithms
- Bergstra and Bengio (2012) Random Search for Hyper-Parameter Optimization
- Bergstra et al. (2011) Algorithms for Hyper-Parameter Optimization
- Turner et al. (2021) Bayesian Optimization is Superior to Random Search for Machine Learning Hyperparameter Tuning: Analysis of the Black-Box Optimization Challenge 2020
- Chen et al. (2018) Bayesian Optimization in AlphaGo
- Golovin et al. (2017) Google Vizier: A Service for Black-Box Optimization
- Thornton et al. (2013) Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms
15.3 Experimental science #
A laboratory experiment satisfies the four conditions almost by definition. It is slow, it consumes material, its outcome cannot be differentiated, and a chemist can vary only a handful of things at once: a catalyst, a solvent, a temperature, a concentration.
Chemistry. Shields et al. (2021) built a framework and an open-source tool for optimizing reaction conditions, and tested it in a way few methods are tested. They measured a large benchmark data set for a palladium-catalyzed reaction, then had 50 expert chemists and engineers play a game in which each chose batches of experiments, receiving the measured yields. Bayesian optimization outperformed human decision-making in both average optimization efficiency, the number of experiments needed, and consistency, the variance of the outcome. Chapter 23 replays that optimization on the published data and looks closely at the comparison with the chemists, who started better and were overtaken (Section 23.4).
Self-driving laboratories. When robots run the experiments and the optimizer chooses the next batch without waiting for a person, the setup is called a self-driving laboratory. Burger et al. (2020) used a mobile robot that "operated autonomously over eight days, performing 688 experiments within a ten-variable experimental space", driven by a batched Bayesian search, and found photocatalyst mixtures six times more active than the initial formulations. Attia et al. (2020) combined Bayesian optimization with a model that predicts a battery's final cycle life from its first few cycles, and identified high-cycle-life fast-charging protocols among 224 candidates in 16 days, compared with over 500 days for exhaustive search without early prediction (Exercise 15.2 asks what that ratio does and does not show).
How large is the gain in general? Adesiji et al. (2026) define the acceleration factor, the number of experiments a reference strategy needs to reach a target divided by the number the optimizer needs. Across 42 studies and 63 benchmarks of self-driving laboratories, the median reported acceleration factor was 6, with a range from 1.3 to 100. The same review found that the benefit after a fixed number of experiments first grows with the number of experiments per dimension and peaks at about 10 to 20 experiments per dimension. Section 36.7 discusses these measures and the role of people in automated experiments, which is mostly to supervise (Kalinin et al., 2024).
Where human experts fit. Experts know things the model does not, and models are patient where experts are not. In a study of semiconductor process development, built as a controlled virtual game, human engineers excelled in the early stages, while the algorithms were far more cost-efficient near the tight tolerances of the target; a strategy with experts first and the algorithm last cut the cost of reaching the target by half compared with experts alone (Kanarik et al., 2023). Expert input is not always a gain: a 2025 preprint documents an industrial case in which adding expert knowledge made Bayesian optimization fail (Weichert et al., 2025).
Sources cited in Section 15.3 7
- Shields et al. (2021) Bayesian reaction optimization as a tool for chemical synthesis
- Burger et al. (2020) A Mobile Robotic Chemist
- Attia et al. (2020) Closed-Loop Optimization of Fast-Charging Protocols for Batteries with Machine Learning
- Adesiji et al. (2026) Benchmarking self-driving labs
- Kalinin et al. (2024) Human-in-the-loop: The future of Machine Learning in Automated Electron Microscopy
- Kanarik et al. (2023) Human–machine collaboration for improving semiconductor process development
- Weichert et al. (2025) When Less is More: A Story of Failing Bayesian Optimization Due to Additional Expert Knowledge
15.4 Robotics and control #
A robot's controller has parameters, gains, timings, and trajectory shapes, that determine how well it walks, flies, or grasps. A simulator can suggest values, but the final tuning happens on the hardware, where each trial takes time, wears the machine, and occasionally breaks it. Trials are noisy, and there are a few to a few dozen parameters.
Gait optimization was an early success. Lizotte et al. (2007) tuned the gait of a quadruped robot for speed and for smoothness with Gaussian process regression and needed dramatically fewer gait evaluations than the local gradient methods then in use. They named the three drawbacks of local methods that a global model removes: they get stuck in local optima, they discard earlier evaluations after each step, and they do not model noise. Calandra et al. (2016) compared automatic gait optimization methods on simulated problems and real robots and concluded that Bayesian optimization is particularly suited to robotics, where good parameters must be found in a small number of experiments.
Bayesian optimization also serves as a component of larger systems. Cully et al. (2015) let a six-legged robot adapt to damage, such as a broken or missing leg, in less than two minutes. Before deployment, the robot built a map of about 13,000 high-performing gaits in simulation; after damage, it ran Bayesian optimization over that map, with the simulated performance as the prior. The search space became a low-dimensional space of behaviors instead of the high-dimensional space of controller parameters, and the authors note that standard Bayesian optimization in the original parameter space did not find working behaviors.
Hardware imposes the constraint of Section 14.4.2: some parameter settings crash the robot. Berkenkamp et al. (2016) applied SafeOpt to tuning a quadrotor's controller, starting from a safe but poorly performing controller and exploring only parameters whose performance stays above a safety threshold with high probability; the tuning ran safely and automatically, without human intervention. For problems with more parameters and larger budgets, trust region methods apply (Section 14.6.2).
When the quality of a robot's behavior is a matter of judgment (does this gait look natural, is this exoskeleton comfortable), the objective is a person's preference, and the methods are those of Part IV; Section 33.6 surveys that work.
Sources cited in Section 15.4 4
- Lizotte et al. (2007) Automatic Gait Optimization with Gaussian Process Regression
- Calandra et al. (2016) Bayesian Optimization for Learning Gaits under Uncertainty
- Cully et al. (2015) Robots That Can Adapt like Animals
- Berkenkamp et al. (2016) Safe Controller Optimization for Quadrotors with Gaussian Processes
15.5 Engineering design #
Bayesian optimization's modern form came from engineering. The efficient global optimization algorithm of Jones et al. (1998) was written for engineering problems in which the number of evaluations is severely limited by time or cost, typically because each one is a long computer simulation, and designing with surrogate models has its own engineering textbook (Forrester et al., 2008). A simulation of the airflow over a wing may take hours; the design may have a dozen shape parameters; and there are constraints, such as a maximum stress, that come out of the same simulation as the objective (Section 14.4). Simulations also come in cheaper and coarser versions, which is the setting of multi-fidelity methods (Section 14.7.2).
Physical machines are tuned the same way. The Linac Coherent Light Source, a free-electron laser, changes configuration several times a day and has to be retuned each time. Duris et al. (2020) tuned groups of its quadrupole magnets with Bayesian optimization, using a Gaussian process whose parameters were fitted from archived scans and whose correlations between magnets came from a simple physical model of the beam; the routine significantly outperformed the facility's existing optimizers. The example shows a pattern common in engineering: the kernel is not a default but encodes what is known about the machine.
Sources cited in Section 15.5 3
- Jones et al. (1998) Efficient Global Optimization of Expensive Black-Box Functions
- Forrester et al. (2008) Engineering Design via Surrogate Modelling: A Practical Guide
- Duris et al. (2020) Bayesian Optimization of a Free-Electron Laser
15.6 Online experiments #
Internet companies tune their products by randomized experiments: some users see variant A, others variant B, and the outcomes are compared. When the variants differ in continuous parameters, such as the weights of a ranking function, the experiment becomes an optimization, and each evaluation is an A/B test that runs for days on live traffic.
This setting stretches the method in three ways. The noise is large, because user behavior varies far more than the differences between variants. The evaluations come in batches, since several variants run at once. And there are constraints: a variant that improves one metric must not degrade others. Letham et al. (2019) developed noisy expected improvement for this setting (Section 14.2.2) and demonstrated it at Facebook on a ranking system and on the flags of a server compiler. The Ax platform packages the approach for general adaptive experimentation (Olson et al., 2025).
Online experiments are also where the scores of Section 13.1 part ways. Every variant is shown to real users, so a bad variant has a cost while it runs: cumulative regret matters, and with a small, fixed set of variants the problem is a bandit (Section 13.2). And the outcomes are several metrics, not one, so someone must say how they trade off. Preference exploration, which learns a decision maker's utility over predicted outcomes from comparisons, was motivated by this problem (Lin et al., 2022). As of September 2026, we found no public report that quantifies preferential Bayesian optimization (PBO) in production A/B tests; Section 34.3 reviews what is known.
Sources cited in Section 15.6 3
- Letham et al. (2019) Constrained Bayesian Optimization with Noisy Experiments
- Olson et al. (2025) Ax: A Platform for Adaptive Experimentation
- Lin et al. (2022) Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes
15.7 People in the loop #
In the applications so far, an instrument produced the number. In a growing set of applications a person is part of the evaluation, in one of two ways.
The person is the system being optimized, and an instrument measures the result. A wearable robot must be tuned to its wearer, and the objective can be physiological. Ding et al. (2018) tuned the peak and offset timing of the hip assistance of a soft exosuit by Bayesian optimization of the metabolic cost of walking. The optimum was found in 21.4 ± 1.0 minutes on average, and metabolic cost fell by 17.4 ± 3.2% compared with walking without the device. The evaluation is slow and noisy: several minutes of walking yield one estimate of energy cost. Kim et al. (2017) had earlier compared Bayesian optimization with a gradient descent method on a simpler task, finding the step frequency that minimizes metabolic cost, and reported faster convergence (12 minutes) with less variability between participants. The person also adapts to the device during the session, so the function being optimized moves; Chapter 24 takes up this problem.
The person is the measuring instrument. Some objectives exist only as judgments. To demonstrate their tuning service, Google engineers optimized a chocolate chip cookie recipe, with parameters that included the amounts of sugar, butter, salt, and cayenne, and the baking time and temperature. Batches were baked, tasters in the company's cafes filled in a survey, and the aggregated ratings went back to the optimizer. "The cookies improved significantly over time", the authors report, and the exercise needed features that real experiments need: marking infeasible recipes (too little butter makes a dough that will not hold together), accepting that a chef changed a suggested recipe, and transferring what was learned at small scale to large-scale baking (Golovin et al., 2017).
Ratings such as these are the simplest way to bring human judgment into the loop, and they have known weaknesses. A rating scale has no fixed zero and no fixed unit; people use it differently from one another and differently at the end of a session than at the beginning (Section 16.1.1). Asking which of two options is better avoids the scale, at the price of less information per answer and a model that no longer has Gaussian observations. That trade is the subject of Part IV, and Chapter 25 works through an example in which the only instrument is the eye.
Three things change when a person is in the loop, whatever the form of the feedback (inference, from the studies cited in this section and in Part VII). The budget is a person's time and patience, tens of evaluations rather than hundreds. The noise is human: it depends on fatigue, on what was shown before, and on how the question is asked. And the person experiences every option tried, so the journey matters as well as the destination, which is the difference between cumulative and simple regret.
Sources cited in Section 15.7 3
- Ding et al. (2018) Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking
- Kim et al. (2017) Human-in-the-Loop Bayesian Optimization of Wearable Device Parameters
- Golovin et al. (2017) Google Vizier: A Service for Black-Box Optimization
15.8 Neighboring methods #
Bayesian optimization shares its machinery with several other fields, and the names are easy to confuse. The differences are mostly differences of goal: what counts as success determines where to evaluate next.
15.8.1 Five neighbors #
Active learning wants an accurate model everywhere, with as few measured inputs as possible. The learner chooses which inputs to have measured (labeled, in the field's vocabulary), and a standard rule is uncertainty sampling (Settles, 2009), which Section 6.5.1 met as the greedy rule for gathering information: ask about the input the model is least sure of, which with a Gaussian process means evaluating where the posterior variance is largest. The machinery is that of Bayesian optimization with the acquisition function changed, and the outcome is different: evaluations spread evenly, including over regions where the function is low and of no interest to an optimizer.
Bayesian experimental design is the general theory behind both. It chooses an experiment to maximize the expected utility of its outcome, where the utility encodes the purpose of the experiment (Chaloner and Verdinelli, 1995). When the purpose is to learn, the utility is the information gained, a criterion that goes back to Lindley (1956) and that MacKay (1992) turned into rules for selecting data (Section 6.4). Bayesian optimization is experimental design with a different utility: the value of the best point found. Entropy search (Section 12.7) makes the link explicit by seeking information about the location of the maximum.
Bandits score every evaluation, not only the final answer (Chapter 13). A bandit algorithm on a continuous domain with a Gaussian process model is Bayesian optimization judged by cumulative regret (Section 13.4), and the upper confidence bound serves both. The fields differ in emphasis: bandit research proves guarantees, usually for many cheap rounds; Bayesian optimization research builds methods for few expensive ones.
Reinforcement learning chooses sequences of actions in an environment whose state changes in response, to maximize reward over time (Sutton and Barto, 2018). A bandit is the special case with a single state: the action affects the reward but not the situation the learner faces next. Bayesian optimization therefore solves a much simpler problem, and it is often used around reinforcement learning rather than instead of it: to tune the hyperparameters of a learning system, as in AlphaGo, or to search directly over the few parameters of a robot's controller, as in gait optimization, where each evaluation runs the whole controller and returns one number.
Evolution strategies search without a model. They keep a population of candidates, or a distribution over candidates, evaluate samples from it, and shift the distribution toward the better samples. CMA-ES, the standard method, adapts a full covariance matrix for its sampling distribution (Hansen and Ostermeier, 2001). Each step is cheap to compute, and nothing limits the number of evaluations, so evolution strategies are the natural choice when evaluations are cheap and plentiful. On a robot control benchmark with a budget of 10,000 evaluations, CMA-ES outperformed every Bayesian optimization method except the trust region method it was compared against (Eriksson et al., 2019). With only tens of evaluations a model-based method has the advantage, because a population needs many evaluations just to estimate a direction (inference; Figure 15.1 shows a small case). One well-known study optimized exoskeleton assistance during walking with an evolution strategy, reducing metabolic energy consumption by 24.2 ± 7.4% compared with no torque (Zhang et al., 2017). When people rate or choose among the candidates of each generation, the method is called interactive evolutionary computation (Takagi, 2001); Section 36.5 compares it with PBO.
15.8.2 Same data, different goals #
The figure below runs four rules on the running objective with the same budget: uncertainty sampling, Bayesian optimization with an upper confidence bound, a simple evolution strategy that keeps one current point and mutates it, and random search. It then scores all four in three ways.
Some things to try.
Start with the model error. Uncertainty sampling ends with the most accurate model: after 30 evaluations its error is about 0.04, against about 0.06 for both Bayesian optimization and random search. Its ticks are spread evenly across the domain. Bayesian optimization's ticks pile up on the tall peak, and its model of the rest of the function stops improving after about fifteen evaluations.
Switch to simple regret. The ranking reverses. Bayesian optimization has found the maximum in essentially every run by 15 evaluations; after 30, uncertainty sampling and random search are still about 0.03 and 0.05 short on average, because they never concentrate on the peak. The evolution strategy is worst here: in many runs it climbs whichever bump is nearest its starting point and stays, since it has no model to tell it that the rest of the domain is unexplored.
Switch to cumulative regret. Bayesian optimization's curve flattens after about ten evaluations, because nearly every later evaluation is near the maximum. Uncertainty sampling and random search keep paying the same amount per evaluation, and end near 22 and 21 against 7. The evolution strategy does better than they do on this score, despite its poor final answer: it spends its evaluations near a good point, although not the best one.
Step back to five evaluations. Early on the rules are hard to tell apart. With little data, the upper confidence bound is dominated by uncertainty, so Bayesian optimization starts as uncertainty sampling and departs from it only once the model has something to exploit (Exercise 15.3).
Active learning, Bayesian optimization, and bandits can share one model and differ in a single line: the rule that turns the posterior into the next evaluation. Before choosing a method, decide what will be scored: the model, the final answer, or everything along the way.
15.8.3 Choosing among them #
Table 15.2 summarizes the comparison. The budgets are orders of magnitude drawn from the examples in this chapter, not rules (inference).
| Method | Goal | Uses a model of the objective | Typical budget | Reach for it when |
|---|---|---|---|---|
| Bayesian optimization | best input found | yes | tens to hundreds | evaluations are expensive and inputs are few |
| Active learning | accurate model everywhere | yes | tens to thousands | the model itself is the product |
| Bayesian experimental design | any stated utility, often information | yes | a few to hundreds | the purpose of the experiment can be written as a utility |
| Bandits | reward summed over all rounds | optional | thousands and more | every evaluation has consequences |
| Reinforcement learning | reward over sequences of actions | optional | very many | actions change the state of the world |
| Evolution strategies | best input found | no | thousands and more | evaluations are cheap, or the dimension is high |
The boundaries are porous. TuRBO borrows the trust region from classical local search (Section 14.6.2); BOHB borrows early stopping from bandits (Section 14.7.2); the knowledge gradient and entropy search are experimental design criteria aimed at the optimum (Section 12.6). What Bayesian optimization contributes to the family is the explicit probabilistic model of an expensive objective, and the habit of asking, before each evaluation, what that evaluation is worth. The next part keeps the model and the habit, and changes the observation: a person's choice between two options instead of a number.
Sources cited in Section 15.8 9
- Settles (2009) Active Learning Literature Survey
- Chaloner and Verdinelli (1995) Bayesian Experimental Design: A Review
- Lindley (1956) On a Measure of the Information Provided by an Experiment
- MacKay (1992) Information-Based Objective Functions for Active Data Selection
- Sutton and Barto (2018) Reinforcement Learning: An Introduction
- Hansen and Ostermeier (2001) Completely Derandomized Self-Adaptation in Evolution Strategies
- Eriksson et al. (2019) Scalable Global Optimization via Local Bayesian Optimization
- Zhang et al. (2017) Human-in-the-loop optimization of exoskeleton assistance during walking
- Takagi (2001) Interactive evolutionary computation: fusion of the capabilities of EC optimization and human evaluation
15.9 Exercises #
For each problem, say which of the methods in Table 15.2 you would try first, and why. (a) A compiler has 30 numeric flags; one benchmark run takes 40 milliseconds. (b) A lab can run one polymer synthesis per day and varies three process settings. (c) A news site wants to choose among five headlines for an article that will be read for the next six hours. (d) An engineering team needs a fast approximation of a slow simulator, accurate over the whole range of four inputs, to use in later studies.
Solution
(a) An evolution strategy or random search: hundreds of thousands of evaluations are affordable, and with evaluations this cheap the overhead of fitting a model would exceed the cost of the evaluations it saves. (b) Bayesian optimization: few inputs, a budget of a few dozen evaluations, each one expensive. (c) A bandit algorithm such as Thompson sampling: there are five arms, every reader shown a worse headline is a loss, and the score is cumulative. (d) Active learning or a space-filling design (Section 11.4): the goal is an accurate model everywhere, not a maximum, so evaluations should go where the model is uncertain.
Attia et al. (2020) report finding good charging protocols in 16 days instead of more than 500. Explain why 500/16 is not the acceleration factor of Bayesian optimization as Adesiji et al. (2026) define it, and what additional comparison would isolate the optimizer's contribution.
Solution
The study changed two things at once: an early-prediction model shortened each experiment, and Bayesian optimization reduced the number of experiments. The 500 days refer to exhaustive search without early prediction, so the ratio of about 31 combines both effects and is measured in time, not in experiments. The acceleration factor compares the number of experiments needed to reach a target with and without the optimizer, everything else equal. To isolate it, one would compare Bayesian optimization with a reference strategy, such as random or exhaustive search, both using early prediction, and count experiments to reach the same cycle life.
The upper confidence bound rule evaluates at the maximizer of (Equation (11.3)). Show that it becomes uncertainty sampling as and pure exploitation at . Use this to explain why Bayesian optimization and uncertainty sampling behave alike during the first few evaluations in Figure 15.1.
Solution
Dividing the score by does not change its maximizer, and as , so the rule picks the input of largest posterior standard deviation, which is uncertainty sampling. At the score is alone, pure exploitation. For a fixed , what matters is the size of the variation in across the domain relative to the variation in . With few evaluations, ranges from near zero at the data to the prior standard deviation elsewhere, while is still close to the constant prior mean over most of the domain, so the uncertainty term decides and the rule explores like uncertainty sampling. As data accumulate, shrinks everywhere and differences in take over.
Sources cited in Section 15.9 2
- Attia et al. (2020) Closed-Loop Optimization of Fast-Charging Protocols for Batteries with Machine Learning
- Adesiji et al. (2026) Benchmarking self-driving labs
Further reading #
- Shahriari et al. (2016) is a survey of Bayesian optimization organized around its applications, with the title's promise of taking the human out of the loop; a good complement to this book, which puts the human back in.
- Shields et al. (2021) is the chemistry study, with the game against 50 chemists; Chapter 23 works through its data.
- Calandra et al. (2016) compares gait optimization methods on real robots and is a careful example of evaluating Bayesian optimization on hardware.
- Letham et al. (2019) describes Bayesian optimization of online experiments, where noise and constraints dominate.
- Golovin et al. (2017) describes a tuning service used at scale, including transfer learning, early stopping, and the cookies.
- Settles (2009) surveys active learning, and Chaloner and Verdinelli (1995) reviews Bayesian experimental design; both are the right starting points for the neighbors of Section 15.8.
- Lattimore and Szepesvári (2020) and Sutton and Barto (2018) are the standard texts on bandits and reinforcement learning.
References
- (2026). Benchmarking self-driving labs. Digital Discovery. Cited in §15.3 §15.9
- (2020). Closed-Loop Optimization of Fast-Charging Protocols for Batteries with Machine Learning. Nature. Cited in §15.1 §15.3 §15.9
- (2012). Random Search for Hyper-Parameter Optimization. Journal of Machine Learning Research. Cited in §15.2
- (2011). Algorithms for Hyper-Parameter Optimization. Advances in Neural Information Processing Systems 24 (NeurIPS 2011). Cited in §15.2
- (2016). Safe Controller Optimization for Quadrotors with Gaussian Processes. IEEE International Conference on Robotics and Automation (ICRA 2016). Cited in §15.1 §15.4
- (2020). A Mobile Robotic Chemist. Nature. Cited in §15.1 §15.3
- (2016). Bayesian Optimization for Learning Gaits under Uncertainty. Annals of Mathematics and Artificial Intelligence. Cited in §15.4
- (1995). Bayesian Experimental Design: A Review. Statistical Science. Cited in §15.8
- (2018). Bayesian Optimization in AlphaGo. preprint Cited in §15.1 §15.2
- (2015). Robots That Can Adapt like Animals. Nature. Cited in §15.4
- (2018). Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking. Science Robotics. Cited in §15.1 §15.7
- (2020). Bayesian Optimization of a Free-Electron Laser. Physical Review Letters. Cited in §15.1 §15.5
- (2019). Scalable Global Optimization via Local Bayesian Optimization. Advances in Neural Information Processing Systems 32 (NeurIPS 2019). Cited in §15.8
- (2008). Engineering Design via Surrogate Modelling: A Practical Guide. Wiley. Cited in §15.5
- (2017). Google Vizier: A Service for Black-Box Optimization. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2017). Cited in §15.1 §15.2 §15.7
- (2001). Completely Derandomized Self-Adaptation in Evolution Strategies. Evolutionary Computation. Cited in §15.8
- (1998). Efficient Global Optimization of Expensive Black-Box Functions. Journal of Global Optimization. Cited in §15.5
- (2024). Human-in-the-loop: The future of Machine Learning in Automated Electron Microscopy. Microscopy Today. doi:10.1093/mictod/qaad096. Cited in §15.3
- (2023). Human–machine collaboration for improving semiconductor process development. Nature. Cited in §15.3
- (2017). Human-in-the-Loop Bayesian Optimization of Wearable Device Parameters. PLOS ONE. Cited in §15.7
- (2020). Bandit Algorithms. Cambridge University Press. doi:10.1017/9781108571401.
- (2019). Constrained Bayesian Optimization with Noisy Experiments. Bayesian Analysis. Cited in §15.1 §15.6
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Cited in §15.6
- (1956). On a Measure of the Information Provided by an Experiment. The Annals of Mathematical Statistics. Cited in §15.8
- (2007). Automatic Gait Optimization with Gaussian Process Regression. Proceedings of the 20th International Joint Conference on Artificial Intelligence (IJCAI 2007). Cited in §15.1 §15.4
- (1992). Information-Based Objective Functions for Active Data Selection. Neural Computation. Cited in §15.8
- (2025). Ax: A Platform for Adaptive Experimentation. International Conference on Automated Machine Learning. Cited in §15.6
- (2009). Active Learning Literature Survey. University of Wisconsin–Madison. non-peer-reviewed Cited in §15.8
- (2016). Taking the Human Out of the Loop: A Review of Bayesian Optimization. Proceedings of the IEEE.
- (2021). Bayesian reaction optimization as a tool for chemical synthesis. Nature. Cited in §15.1 §15.3
- (2012). Practical Bayesian Optimization of Machine Learning Algorithms. Advances in Neural Information Processing Systems 25 (NeurIPS 2012). Cited in §15.2
- (2018). Reinforcement Learning: An Introduction. MIT Press. Cited in §15.8
- (2001). Interactive evolutionary computation: fusion of the capabilities of EC optimization and human evaluation. Proceedings of the IEEE. Cited in §15.8
- (2013). Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms. Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2013). Cited in §15.2
- (2021). Bayesian Optimization is Superior to Random Search for Machine Learning Hyperparameter Tuning: Analysis of the Black-Box Optimization Challenge 2020. NeurIPS 2020 Competition and Demonstration Track. Cited in §15.2
- (2025). When Less is More: A Story of Failing Bayesian Optimization Due to Additional Expert Knowledge. arXiv. preprint Cited in §15.3
- (2017). Human-in-the-loop optimization of exoskeleton assistance during walking. Science. Cited in §15.8