高斯过程
高斯过程是函数上的概率分布。借助它,可以在获得任何数据之前表明哪些函数是合理的,并在数据到来时精确地更新这一信念。几乎所有贝叶斯优化都以高斯过程为代理模型;在偏好贝叶斯优化中,它用来为一个人隐藏的效用建模。
前三章沿着这一信念的发展过程展开。第一章让贝叶斯线性回归的特征不断增多,由此构造先验,读者可以从中抽取函数。第二章以观测为条件更新先验,并解读所得结果:后验均值在数据之间插值,后验不确定性在已观测之处收缩。第三章讨论如何选择核函数及其超参数,并说明为什么在 2 维中好用的默认设置,到 50 维时会悄然失效。第四章深入模型内部,考察核函数定义的函数空间、核函数的特征值与频谱,以及高斯过程作为随机对象的含义:这些分析是第三部分中理论保证的依据。
阅读本部分需要第 4.5 节中的高斯条件化公式,以及第 5 章中后验的概念。
本部分各章
- 7 函数上的分布
从模型权重上的先验过渡到整个函数上的先验:特征、由特征诱导的核函数、高斯过程的定义,以及核函数与超参数变化时,从高斯过程中抽取的函数呈现怎样的形态。
- 8 高斯过程回归
以观测为条件更新高斯过程:预测方程、后验均值与不确定性对数据和噪声的响应,以及稳定的计算方法。
- 9 核函数与超参数
核函数、长度尺度与噪声水平的选择:核函数族及其组合方式;为每个输入设置单独的长度尺度;边际似然作为数据对模型的裁决;超参数的拟合方法及其失效方式;长度尺度先验为何必须随输入个数增长;以及如何检验结果。
- 10 核函数背后的分析
遗憾界背后的数学:把函数看作向量;核函数定义的再生核 Hilbert 空间,以及其范数界的含义;Mercer 特征展开;Bochner 定理给出的谱观点与随机特征;特征值衰减如何决定信息增益;以及高斯过程作为随机对象的精确含义。
第二部分参考文献
本部分各章共引用 42 篇文献。
- (1950). Theory of Reproducing Kernels. Transactions of the American Mathematical Society. 第 10 章
- (2020). BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). 第 8 章
- (2018). Approximation Beats Concentration? An Approximation View on Inference with Smooth Radial Kernels. Proceedings of the 31st Conference on Learning Theory. 第 10 章
- (2004). Reproducing Kernel Hilbert Spaces in Probability and Statistics. Springer. 第 10 章
- (1933). Monotone Funktionen, Stieltjessche Integrale und harmonische Analyse. Mathematische Annalen. 第 10 章
- (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. 第 10 章
- (2006). Elements of Information Theory. Wiley. 第 10 章
- (2026). Sample Path Regularity of Gaussian Processes from the Covariance Kernel. Analysis and Applications. 第 10 章
- (1973). The Reproducing Kernel Hilbert Space Structure of the Sample Paths of a Gaussian Process. Zeitschrift für Wahrscheinlichkeitstheorie und Verwandte Gebiete. 第 10 章
- (2013). Structure Discovery in Nonparametric Regression through Compositional Kernel Search. Proceedings of the 30th International Conference on Machine Learning (ICML 2013). 第 9 章
- (2023). Bayesian Optimization. Cambridge University Press. 第 7 章 第 8 章 第 9 章
- (2019). A Visual Exploration of Gaussian Processes. Distill. doi:10.23915/distill.00017. 第 7 章 第 8 章
- (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. 第 7 章 第 9 章
- (2018). Gaussian Processes and Kernel Methods: A Review on Connections and Equivalences. arXiv preprint. 预印本 第 8 章 第 10 章
- (1970). A Correspondence Between Bayesian Estimation on Stochastic Processes and Smoothing by Splines. The Annals of Mathematical Statistics. 第 10 章
- (1971). Some Results on Tchebycheffian Spline Functions. Journal of Mathematical Analysis and Applications. 第 10 章
- (1933). Grundbegriffe der Wahrscheinlichkeitsrechnung. Springer. 第 10 章
- (1951). A Statistical Approach to Some Basic Mine Valuation Problems on the Witwatersrand. Journal of the Southern African Institute of Mining and Metallurgy. 第 8 章
- (2001). Stochastic Processes with Sample Paths in Reproducing Kernel Hilbert Spaces. Transactions of the American Mathematical Society. 第 10 章
- (1963). Principles of Geostatistics. Economic Geology. 第 8 章
- (1909). Functions of Positive and Negative Type, and Their Connection with the Theory of Integral Equations. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character. 第 10 章
- (2026e). BoTorch CHANGELOG. GitHub. 软件 第 9 章
- (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. 软件 第 9 章
- (2026k). botorch/models/utils/gpytorch_modules.py. GitHub. 软件 第 9 章
- (1996). Bayesian Learning for Neural Networks. Springer. 第 7 章 第 9 章
- (2025b). Understanding High-Dimensional Bayesian Optimization. ICML 2025, PMLR 267:47902-47923. 第 9 章
- (2012). The Matrix Cookbook. Technical University of Denmark. 非同行评审 第 9 章
- (2005). A Unifying View of Sparse Approximate Gaussian Process Regression. Journal of Machine Learning Research. 第 8 章
- (2007). Random Features for Large-Scale Kernel Machines. Advances in Neural Information Processing Systems 20 (NeurIPS 2007). 第 8 章 第 10 章
- (2006). Gaussian Processes for Machine Learning. MIT Press. 第 7 章 第 8 章 第 9 章 第 10 章
- (2016). Approximation of Eigenfunctions in Kernel-Based Spaces. Advances in Computational Mathematics. 第 10 章
- (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. 第 10 章
- (2012). Practical Bayesian Optimization of Machine Learning Algorithms. Advances in Neural Information Processing Systems 25 (NeurIPS 2012). 第 7 章 第 9 章
- (2010). Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. ICML 2010. 第 10 章
- (1999). Interpolation of Spatial Data: Some Theory for Kriging. Springer. 第 7 章 第 9 章
- (2008). Support Vector Machines. Springer. 第 10 章
- (2009). Variational Learning of Inducing Variables in Sparse Gaussian Processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics (AISTATS 2009). 第 8 章
- (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. 第 10 章
- (2004). Scattered Data Approximation. Cambridge University Press. 第 10 章
- (1996). Gaussian Processes for Regression. Advances in Neural Information Processing Systems 8 (NeurIPS 1995). 第 8 章
- (2020). Efficiently Sampling Functions from Gaussian Process Posteriors. Proceedings of the 37th International Conference on Machine Learning (ICML 2020). 第 8 章 第 10 章
- (2025b). Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization. ICLR 2025 (oral). 第 9 章