贝叶斯优化
第一部分:基础
EN

基础

贝叶斯优化由少数几件数学工具构成。这一部分面向会编写软件、但离校后未再使用概率的读者,从基础开始逐一构建这些工具。概率为不确定性赋予数值。线性代数提供同时计算许多不确定数值的方法。在本书所需的全部运算下都保持封闭的分布族,只有高斯分布这一族,因此以数据为条件这一步(也是后续各章的核心步骤)可以精确求解。贝叶斯推断把这一步转化为学习,信息论则衡量单次观测的价值。

每章引入一件工具,并立即将其投入使用。读完这一部分,读者将亲手推导出(而不只是读到)一个公式,下一部分会把它发展为高斯过程回归。

熟悉概率与线性代数的读者可以略读这一部分,待后续章节引用时,再回到相应小节。

本部分各章

  1. 1 优化写不出公式的函数

    目标函数怎样才算黑箱,为什么每次评估都很宝贵,为什么出路在于为目标函数建模,并把每次评估用在能学到最多的地方。附全书地图。

  2. 2 概率:为不确定性记账

    把概率理解为分配给各种可能性的信念预算:随机变量,离散分布与连续分布,以及推出其余一切的两条规则,即加法规则与乘法规则。贝叶斯定理由这两条规则一行推出,期望、方差与独立性则补全了这套工具。

  3. 3 不确定性的线性代数

    向量、作为空间映射的矩阵、正定矩阵、特征向量、Cholesky 分解、行列式与分块矩阵:高斯过程所需的线性代数。每个概念都配有一幅图,并借助一个三维协方差矩阵揭示二维图中看不到的现象。

  4. 4 高斯分布

    全书赖以运转的分布:它在一维和多维中的形状;高斯分布经线性映射后为何仍是高斯分布,以及如何由此得到采样方法;借助 Schur 补推导出的条件化公式,高斯过程回归直接沿用这一公式。

  5. 5 贝叶斯推断

    先以硬币、再以直线和平面为例,精确求出先验、似然、后验与预测分布。说明点估计为何无法指示下一步的评估位置,证据如何权衡不同模型,以及直线权重的后验为何是整个函数上的后验的铺垫。

  6. 6 度量信息

    从单个结果的意外度出发,依次建立熵、KL 散度与互信息;介绍 Lindley 用于选择下一个问题的期望信息增益;最后讨论高斯过程的信息增益,其最大值决定书中所有的遗憾界,并随输入维度迅速增长。

第一部分参考文献

本部分各章共引用 53 篇文献。

  1. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). 第 5 章
  2. Balandat, M., Karrer, B., Jiang, D. R., Daulton, S., Letham, B., Wilson, A. G., and Bakshy, E. (2020). BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). 第 2 章
  3. Bar-Hillel, M. (1980). The Base-Rate Fallacy in Probability Judgments. Acta Psychologica. 第 2 章
  4. Bayes, T. (1763). An Essay towards Solving a Problem in the Doctrine of Chances. Philosophical Transactions of the Royal Society of London. 第 2 章
  5. Bergstra, J., and Bengio, Y. (2012). Random Search for Hyper-Parameter Optimization. Journal of Machine Learning Research. 第 1 章
  6. Bishop, C. M. (2006). Pattern Recognition and Machine Learning. Springer. 第 2 章 第 4 章 第 5 章 第 6 章
  7. Blitzstein, J. K., and Hwang, J. (2019). Introduction to Probability. Chapman and Hall/CRC. 第 2 章 第 4 章
  8. Box, G. E. P., and Muller, M. E. (1958). A Note on the Generation of Random Normal Deviates. The Annals of Mathematical Statistics. 第 4 章
  9. Chaloner, K., and Verdinelli, I. (1995). Bayesian Experimental Design: A Review. Statistical Science. 第 6 章
  10. Chu, W., and Ghahramani, Z. (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. 第 5 章
  11. Cover, T. M., and Thomas, J. A. (2006). Elements of Information Theory. Wiley. 第 4 章 第 6 章
  12. Cox, R. T. (1946). Probability, Frequency and Reasonable Expectation. American Journal of Physics. 第 2 章
  13. Ding, Y., Kim, M., Kuindersma, S., and Walsh, C. J. (2018). Human-in-the-Loop Optimization of Hip Assistance with a Soft Exosuit during Walking. Science Robotics. 第 2 章
  14. Doumont, C., Fan, D., Maus, N., Gardner, J. R., Moss, H., and Pleiss, G. (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). 第 5 章
  15. Frazier, P. I. (2018). A Tutorial on Bayesian Optimization. arXiv. 预印本 第 1 章
  16. Gabry, J., Simpson, D., Vehtari, A., Betancourt, M., and Gelman, A. (2019). Visualization in Bayesian Workflow. Journal of the Royal Statistical Society Series A: Statistics in Society. 第 5 章
  17. Galton, F. (1886). Regression Towards Mediocrity in Hereditary Stature. The Journal of the Anthropological Institute of Great Britain and Ireland. 第 4 章
  18. Gardner, J. R., Pleiss, G., Bindel, D., Weinberger, K. Q., and Wilson, A. G. (2018). GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration. Advances in Neural Information Processing Systems 31 (NeurIPS 2018). 第 3 章
  19. Garnett, R. (2023). Bayesian Optimization. Cambridge University Press. 第 1 章
  20. Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B. (2013). Bayesian Data Analysis. Chapman and Hall/CRC. 第 5 章
  21. Gigerenzer, G., and Hoffrage, U. (1995). How to Improve Bayesian Reasoning Without Instruction: Frequency Formats. Psychological Review. 第 2 章
  22. Golub, G. H., and Van Loan, C. F. (2013). Matrix Computations. Johns Hopkins University Press. 第 3 章
  23. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. 第 1 章
  24. Hennig, P., and Schuler, C. J. (2012). Entropy Search for Information-Efficient Global Optimization. Journal of Machine Learning Research. 第 6 章
  25. Hernández-Lobato, J. M., Hoffman, M. W., and Ghahramani, Z. (2014). Predictive Entropy Search for Efficient Global Optimization of Black-box Functions. Advances in Neural Information Processing Systems 27 (NeurIPS 2014). 第 6 章
  26. Higham, N. J. (2002). Computing the Nearest Correlation Matrix: A Problem from Finance. IMA Journal of Numerical Analysis. 第 3 章
  27. Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M. (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. 预印本 第 6 章
  28. Hvarfner, C., Hellsten, E. O., and Nardi, L. (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. 第 6 章
  29. Jaynes, E. T. (2003). Probability Theory: The Logic of Science. Cambridge University Press. 第 2 章
  30. Jones, D. R., Schonlau, M., and Welch, W. J. (1998). Efficient Global Optimization of Expensive Black-Box Functions. Journal of Global Optimization. 第 1 章
  31. Kahneman, D., and Tversky, A. (1973). On the Psychology of Prediction. Psychological Review. 第 2 章
  32. Kontsevich, L. L., and Tyler, C. W. (1999). Bayesian Adaptive Estimation of Psychometric Slope and Threshold. Vision Research. 第 6 章
  33. Kullback, S., and Leibler, R. A. (1951). On Information and Sufficiency. The Annals of Mathematical Statistics. 第 6 章
  34. Kushner, H. J. (1964). A New Method of Locating the Maximum Point of an Arbitrary Multipeak Curve in the Presence of Noise. Journal of Basic Engineering. 第 1 章
  35. Lindley, D. V. (1956). On a Measure of the Information Provided by an Experiment. The Annals of Mathematical Statistics. 第 6 章
  36. MacKay, D. J. C. (1992). Information-Based Objective Functions for Active Data Selection. Neural Computation. 第 6 章
  37. MacKay, D. J. C. (2003). Information Theory, Inference, and Learning Algorithms. Cambridge University Press. 第 2 章 第 5 章 第 6 章
  38. Močkus, J. (1975). On Bayesian Methods for Seeking the Extremum. Optimization Techniques IFIP Technical Conference. 第 1 章
  39. Murphy, K. P. (2022). Probabilistic Machine Learning: An Introduction. MIT Press. 第 4 章
  40. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems. 第 6 章
  41. Petersen, K. B., and Pedersen, M. S. (2012). The Matrix Cookbook. Technical University of Denmark. 非同行评审 第 3 章 第 4 章
  42. Rasmussen, C. E., and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press. 第 3 章 第 4 章 第 5 章
  43. Sanderson, G. (2016). Essence of Linear Algebra. Video series, 3Blue1Brown. 非同行评审 第 3 章
  44. Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and de Freitas, N. (2016). Taking the Human Out of the Loop: A Review of Bayesian Optimization. Proceedings of the IEEE. 第 1 章
  45. Shannon, C. E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal. 第 6 章
  46. Snoek, J., Larochelle, H., and Adams, R. P. (2012). Practical Bayesian Optimization of Machine Learning Algorithms. Advances in Neural Information Processing Systems 25 (NeurIPS 2012). 第 1 章 第 3 章 第 4 章
  47. Snoek, J., Rippel, O., Swersky, K., Kiros, R., Satish, N., Sundaram, N., … Adams, R. P. (2015). Scalable Bayesian Optimization Using Deep Neural Networks. Proceedings of the 32nd International Conference on Machine Learning (ICML 2015). 第 5 章
  48. Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M. (2010). Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. ICML 2010. 第 6 章
  49. Strang, G. (2016). Introduction to Linear Algebra. Wellesley-Cambridge Press. 第 3 章
  50. Thompson, W. R. (1933). On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika. 第 5 章
  51. Vakili, S., Khezeli, K., and Picheny, V. (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. 第 6 章
  52. Wang, Z., and Jegelka, S. (2017). Max-value Entropy Search for Efficient Bayesian Optimization. Proceedings of the 34th International Conference on Machine Learning (ICML 2017). 第 6 章
  53. Watson, A. B., and Pelli, D. G. (1983). QUEST: A Bayesian Adaptive Psychometric Method. Perception & Psychophysics. 第 6 章