Bayesian Optimization
Part II: Gaussian Processes
中文

Gaussian Processes

A Gaussian process is a probability distribution over functions: a way of saying, before any data arrive, which functions you find plausible, and of updating that belief exactly when data do arrive. It is the surrogate model in almost all Bayesian optimization, and in preferential Bayesian optimization it becomes the model of a person's hidden utility.

The first three chapters follow the life of that belief. The first builds the prior from Bayesian linear regression with ever more features, and lets you draw functions from it. The second conditions it on observations and reads the result: a posterior mean that interpolates and a posterior uncertainty that collapses where you have looked. The third asks how to choose the kernel and its hyperparameters, and why a default that works in two dimensions can quietly fail in fifty. The fourth looks under the model, at the space of functions a kernel defines, its eigenvalues and spectrum, and what a Gaussian process is as a random object: the analysis that the guarantees of Part III rely on.

The part needs the Gaussian conditioning formula from Section 4.5 and the idea of a posterior from Chapter 5.

Chapters in this part

  1. 7 Distributions over Functions

    From a prior over a model's weights to a prior over whole functions: features, the kernel they induce, the definition of a Gaussian process, and what functions drawn from it look like as the kernel and its hyperparameters change.

  2. 8 Gaussian Process Regression

    Conditioning a Gaussian process on observations: the predictive equations, how the posterior mean and uncertainty respond to data and noise, and how to compute them stably.

  3. 9 Kernels and Hyperparameters

    Which kernel, which lengthscales, how much noise: the kernel family and how kernels combine, one lengthscale per input, the marginal likelihood as the data's verdict on a model, how hyperparameters are fitted and how the fit fails, why the lengthscale prior must grow with the number of inputs, and how to check the result.

  4. 10 The Analysis Behind Kernels

    The mathematics under the regret bounds: functions as vectors, the reproducing kernel Hilbert space a kernel defines and what a bound on its norm means, Mercer's eigen-expansion, Bochner's spectral view and random features, how eigenvalue decay sets the information gain, and what a Gaussian process is as a random object.

References for Part II

42 works cited across this part's chapters.

  1. Aronszajn, N. (1950). Theory of Reproducing Kernels. Transactions of the American Mathematical Society. Ch. 10
  2. Balandat, M., Karrer, B., Jiang, D. R., Daulton, S., Letham, B., Wilson, A. G., and Bakshy, E. (2020). BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). Ch. 8
  3. Belkin, M. (2018). Approximation Beats Concentration? An Approximation View on Inference with Smooth Radial Kernels. Proceedings of the 31st Conference on Learning Theory. Ch. 10
  4. Berlinet, A., and Thomas-Agnan, C. (2004). Reproducing Kernel Hilbert Spaces in Probability and Statistics. Springer. Ch. 10
  5. Bochner, S. (1933). Monotone Funktionen, Stieltjessche Integrale und harmonische Analyse. Mathematische Annalen. Ch. 10
  6. Chowdhury, S. R., and Gopalan, A. (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. Ch. 10
  7. Cover, T. M., and Thomas, J. A. (2006). Elements of Information Theory. Wiley. Ch. 10
  8. Da Costa, N., Pförtner, M., Da Costa, L., and Hennig, P. (2026). Sample Path Regularity of Gaussian Processes from the Covariance Kernel. Analysis and Applications. Ch. 10
  9. Driscoll, M. F. (1973). The Reproducing Kernel Hilbert Space Structure of the Sample Paths of a Gaussian Process. Zeitschrift für Wahrscheinlichkeitstheorie und Verwandte Gebiete. Ch. 10
  10. Duvenaud, D., Lloyd, J., Grosse, R., Tenenbaum, J., and Ghahramani, Z. (2013). Structure Discovery in Nonparametric Regression through Compositional Kernel Search. Proceedings of the 30th International Conference on Machine Learning (ICML 2013). Ch. 9
  11. Garnett, R. (2023). Bayesian Optimization. Cambridge University Press. Ch. 7 Ch. 8 Ch. 9
  12. Görtler, J., Kehlbeck, R., and Deussen, O. (2019). A Visual Exploration of Gaussian Processes. Distill. doi:10.23915/distill.00017. Ch. 7 Ch. 8
  13. Hvarfner, C., Hellsten, E. O., and Nardi, L. (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. Ch. 7 Ch. 9
  14. Kanagawa, M., Hennig, P., Sejdinovic, D., and Sriperumbudur, B. K. (2018). Gaussian Processes and Kernel Methods: A Review on Connections and Equivalences. arXiv preprint. preprint Ch. 8 Ch. 10
  15. Kimeldorf, G. S., and Wahba, G. (1970). A Correspondence Between Bayesian Estimation on Stochastic Processes and Smoothing by Splines. The Annals of Mathematical Statistics. Ch. 10
  16. Kimeldorf, G., and Wahba, G. (1971). Some Results on Tchebycheffian Spline Functions. Journal of Mathematical Analysis and Applications. Ch. 10
  17. Kolmogoroff, A. (1933). Grundbegriffe der Wahrscheinlichkeitsrechnung. Springer. Ch. 10
  18. Krige, D. G. (1951). A Statistical Approach to Some Basic Mine Valuation Problems on the Witwatersrand. Journal of the Southern African Institute of Mining and Metallurgy. Ch. 8
  19. Lukić, M. N., and Beder, J. H. (2001). Stochastic Processes with Sample Paths in Reproducing Kernel Hilbert Spaces. Transactions of the American Mathematical Society. Ch. 10
  20. Matheron, G. (1963). Principles of Geostatistics. Economic Geology. Ch. 8
  21. Mercer, J. (1909). Functions of Positive and Negative Type, and Their Connection with the Theory of Integral Equations. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character. Ch. 10
  22. Meta Platforms, Inc. (2026e). BoTorch CHANGELOG. GitHub. software Ch. 9
  23. Meta Platforms, Inc. (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Ch. 9
  24. Meta Platforms, Inc. (2026k). botorch/models/utils/gpytorch_modules.py. GitHub. software Ch. 9
  25. Neal, R. M. (1996). Bayesian Learning for Neural Networks. Springer. Ch. 7 Ch. 9
  26. Papenmeier, L., Poloczek, M., and Nardi, L. (2025b). Understanding High-Dimensional Bayesian Optimization. ICML 2025, PMLR 267:47902-47923. Ch. 9
  27. Petersen, K. B., and Pedersen, M. S. (2012). The Matrix Cookbook. Technical University of Denmark. non-peer-reviewed Ch. 9
  28. Quiñonero-Candela, J., and Rasmussen, C. E. (2005). A Unifying View of Sparse Approximate Gaussian Process Regression. Journal of Machine Learning Research. Ch. 8
  29. Rahimi, A., and Recht, B. (2007). Random Features for Large-Scale Kernel Machines. Advances in Neural Information Processing Systems 20 (NeurIPS 2007). Ch. 8 Ch. 10
  30. Rasmussen, C. E., and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press. Ch. 7 Ch. 8 Ch. 9 Ch. 10
  31. Santin, G., and Schaback, R. (2016). Approximation of Eigenfunctions in Kernel-Based Spaces. Advances in Computational Mathematics. Ch. 10
  32. Scarlett, J., Bogunovic, I., and Cevher, V. (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. Ch. 10
  33. Snoek, J., Larochelle, H., and Adams, R. P. (2012). Practical Bayesian Optimization of Machine Learning Algorithms. Advances in Neural Information Processing Systems 25 (NeurIPS 2012). Ch. 7 Ch. 9
  34. Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M. (2010). Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. ICML 2010. Ch. 10
  35. Stein, M. L. (1999). Interpolation of Spatial Data: Some Theory for Kriging. Springer. Ch. 7 Ch. 9
  36. Steinwart, I., and Christmann, A. (2008). Support Vector Machines. Springer. Ch. 10
  37. Titsias, M. (2009). Variational Learning of Inducing Variables in Sparse Gaussian Processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics (AISTATS 2009). Ch. 8
  38. Vakili, S., Khezeli, K., and Picheny, V. (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 10
  39. Wendland, H. (2004). Scattered Data Approximation. Cambridge University Press. Ch. 10
  40. Williams, C. K. I., and Rasmussen, C. E. (1996). Gaussian Processes for Regression. Advances in Neural Information Processing Systems 8 (NeurIPS 1995). Ch. 8
  41. Wilson, J. T., Borovitskiy, V., Terenin, A., Mostowski, P., and Deisenroth, M. P. (2020). Efficiently Sampling Functions from Gaussian Process Posteriors. Proceedings of the 37th International Conference on Machine Learning (ICML 2020). Ch. 8 Ch. 10
  42. Xu, Z., Wang, H., Phillips, J. M., and Zhe, S. (2025b). Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization. ICLR 2025 (oral). Ch. 9