Foundations
Bayesian optimization is built from a small number of mathematical tools, and this part builds each of them from the ground up for a reader who writes software but has not used probability since school. Probability gives uncertainty a number. Linear algebra gives a way to compute with many uncertain numbers at once. The Gaussian distribution is the one family of distributions that stays closed under every operation the book needs, so that conditioning on data, the central step of every later chapter, has an exact answer. Bayesian inference turns that step into learning, and information theory measures how much a single observation is worth.
Each chapter introduces one tool and immediately puts it to work. By the end of the part you will have derived, not just read, the formula that the next part turns into Gaussian process regression.
Readers fluent in probability and linear algebra can skim this part and return to individual sections when a later chapter points back to them.
Chapters in this part
- 1 Optimizing What You Cannot Write Down
What makes an objective a black box, why each evaluation is precious, and why the answer is to model the objective and spend every evaluation where it teaches the most. A map of the book.
- 2 Probability as Bookkeeping for Uncertainty
Probability as a budget of belief spread over possibilities: random variables, discrete and continuous distributions, and the two rules everything else follows from, the sum rule and the product rule. Bayes' rule falls out of them in one line, and expectation, variance, and independence complete the toolkit.
- 3 The Linear Algebra of Uncertainty
Vectors, matrices as maps of space, positive definite matrices, eigenvectors, the Cholesky factorization, determinants, and block matrices: the linear algebra a Gaussian process needs, with a picture for each idea and a covariance matrix in three dimensions to show what two dimensions hide.
- 4 The Gaussian Distribution
The one distribution the whole book runs on: its shape in one and many dimensions, why a linear map of a Gaussian is Gaussian and how that gives a sampler, and the conditioning formula, derived through the Schur complement, that Gaussian process regression applies unchanged.
- 5 Bayesian Inference
Prior, likelihood, posterior, and predictive distribution, worked out exactly for a coin and then for a line and a plane. Why a point estimate cannot say where to look next, how the evidence weighs models, and why the posterior over a line's weights is the stepping stone to a posterior over whole functions.
- 6 Measuring Information
Entropy, KL divergence, and mutual information, built from the surprise of a single outcome; Lindley's expected information gain for choosing what to ask next; and the information gain of a Gaussian process, whose maximum sets every regret bound in the book and grows quickly with the input dimension.
References for Part I
53 works cited across this part's chapters.
- (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Ch. 5
- (2020). BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). Ch. 2
- (1980). The Base-Rate Fallacy in Probability Judgments. Acta Psychologica. Ch. 2
- (1763). An Essay towards Solving a Problem in the Doctrine of Chances. Philosophical Transactions of the Royal Society of London. Ch. 2
- (2012). Random Search for Hyper-Parameter Optimization. Journal of Machine Learning Research. Ch. 1
- (2006). Pattern Recognition and Machine Learning. Springer. Ch. 2 Ch. 4 Ch. 5 Ch. 6
- (2019). Introduction to Probability. Chapman and Hall/CRC. Ch. 2 Ch. 4
- (1958). A Note on the Generation of Random Normal Deviates. The Annals of Mathematical Statistics. Ch. 4
- (1995). Bayesian Experimental Design: A Review. Statistical Science. Ch. 6
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Ch. 5
- (2006). Elements of Information Theory. Wiley. Ch. 4 Ch. 6
- (1946). Probability, Frequency and Reasonable Expectation. American Journal of Physics. Ch. 2
- (2018). Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking. Science Robotics. Ch. 2
- (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). Ch. 5
- (2018). A Tutorial on Bayesian Optimization. arXiv. preprint Ch. 1
- (2019). Visualization in Bayesian Workflow. Journal of the Royal Statistical Society Series A: Statistics in Society. Ch. 5
- (1886). Regression Towards Mediocrity in Hereditary Stature. The Journal of the Anthropological Institute of Great Britain and Ireland. Ch. 4
- (2018). GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration. Advances in Neural Information Processing Systems 31 (NeurIPS 2018). Ch. 3
- (2023). Bayesian Optimization. Cambridge University Press. Ch. 1
- (2013). Bayesian Data Analysis. Chapman and Hall/CRC. Ch. 5
- (1995). How to Improve Bayesian Reasoning Without Instruction: Frequency Formats. Psychological Review. Ch. 2
- (2013). Matrix Computations. Johns Hopkins University Press. Ch. 3
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 1
- (2012). Entropy Search for Information-Efficient Global Optimization. Journal of Machine Learning Research. Ch. 6
- (2014). Predictive Entropy Search for Efficient Global Optimization of Black-box Functions. Advances in Neural Information Processing Systems 27 (NeurIPS 2014). Ch. 6
- (2002). Computing the Nearest Correlation Matrix: A Problem from Finance. IMA Journal of Numerical Analysis. Ch. 3
- (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Ch. 6
- (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. Ch. 6
- (2003). Probability Theory: The Logic of Science. Cambridge University Press. Ch. 2
- (1998). Efficient Global Optimization of Expensive Black-Box Functions. Journal of Global Optimization. Ch. 1
- (1973). On the Psychology of Prediction. Psychological Review. Ch. 2
- (1999). Bayesian Adaptive Estimation of Psychometric Slope and Threshold. Vision Research. Ch. 6
- (1951). On Information and Sufficiency. The Annals of Mathematical Statistics. Ch. 6
- (1964). A New Method of Locating the Maximum Point of an Arbitrary Multipeak Curve in the Presence of Noise. Journal of Basic Engineering. Ch. 1
- (1956). On a Measure of the Information Provided by an Experiment. The Annals of Mathematical Statistics. Ch. 6
- (1992). Information-Based Objective Functions for Active Data Selection. Neural Computation. Ch. 6
- (2003). Information Theory, Inference, and Learning Algorithms. Cambridge University Press. Ch. 2 Ch. 5 Ch. 6
- (1975). On Bayesian Methods for Seeking the Extremum. Optimization Techniques IFIP Technical Conference. Ch. 1
- (2022). Probabilistic Machine Learning: An Introduction. MIT Press. Ch. 4
- (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems. Ch. 6
- (2012). The Matrix Cookbook. Technical University of Denmark. non-peer-reviewed Ch. 3 Ch. 4
- (2006). Gaussian Processes for Machine Learning. MIT Press. Ch. 3 Ch. 4 Ch. 5
- (2016). Essence of Linear Algebra. Video series, 3Blue1Brown. non-peer-reviewed Ch. 3
- (2016). Taking the Human Out of the Loop: A Review of Bayesian Optimization. Proceedings of the IEEE. Ch. 1
- (1948). A Mathematical Theory of Communication. Bell System Technical Journal. Ch. 6
- (2012). Practical Bayesian Optimization of Machine Learning Algorithms. Advances in Neural Information Processing Systems 25 (NeurIPS 2012). Ch. 1 Ch. 3 Ch. 4
- (2015). Scalable Bayesian Optimization Using Deep Neural Networks. Proceedings of the 32nd International Conference on Machine Learning (ICML 2015). Ch. 5
- (2010). Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. ICML 2010. Ch. 6
- (2016). Introduction to Linear Algebra. Wellesley-Cambridge Press. Ch. 3
- (1933). On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika. Ch. 5
- (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 6
- (2017). Max-value Entropy Search for Efficient Bayesian Optimization. Proceedings of the 34th International Conference on Machine Learning (ICML 2017). Ch. 6
- (1983). QUEST: A Bayesian Adaptive Psychometric Method. Perception & Psychophysics. Ch. 6