Notation
The book uses one notation throughout, and this appendix collects it. Each entry points to the section that introduces the symbol, where it is explained in words before it is used in a formula. Where the literature uses several conventions, the entry says which one the book follows.
A few typographic rules hold everywhere. Scalars are italic (, ), vectors are bold lowercase (, ), and matrices are bold uppercase (, ). A transpose is written . Inputs live in a domain , usually the unit cube after rescaling, where is the number of inputs; when an index runs over the inputs, as for one lengthscale per input, it is . A finite set of candidate inputs is also written , with elements, even where the papers the book reports write . "Larger is better" throughout: the book maximizes, and a function that is naturally minimized, such as an error rate, is negated.
A.1 Sets, vectors, and matrices #
| Symbol | Meaning | Introduced |
|---|---|---|
| , | the real numbers; vectors of real numbers | Section 3.1 |
| the number of inputs (dimension of ) | Section 1.2, Section 3.1 | |
| , | a vector and its -th entry | Section 3.1 |
| , | transpose of a vector and of a matrix | Section 3.1, Section 3.2.3 |
| identity matrix | Section 3.2.2 | |
| inverse of (computed by solving, never formed) | Section 3.2.2 | |
| lower-triangular Cholesky factor, | Section 3.5.3 | |
| the solution of | Section 8.4 | |
| , | determinant, and its logarithm (from the Cholesky diagonal) | Section 3.6, Section 3.6.1 |
| trace, the sum of the diagonal | Section 6.2.2, Section 9.4.1 |
A.2 Probability #
| Symbol | Meaning | Introduced |
|---|---|---|
| probability of an event | Section 2.1.2 | |
| , | density or mass function; conditional on | Section 2.1.3 |
| is distributed according to | Section 2.1.3 | |
| the observed data | Section 2.5 | |
| , | expectation; expectation under the posterior after observations | Section 2.6.1, Section 12.1 |
| , | variance; covariance | Section 2.6.2, Section 2.6.3 |
| , | Gaussian with mean and variance; with mean vector and covariance matrix | Section 4.1, Section 4.2 |
| , | standard normal density and cumulative distribution function | Section 4.1.1 |
| logistic function , the Bradley-Terry link | Section 16.4 | |
| , | entropy of a distribution or of a random variable | Section 6.1.2 |
| Kullback-Leibler divergence | Section 6.2.1 | |
| mutual information | Section 6.3 | |
| failure probability: a high-probability statement holds with probability at least | Section 13.2.2 | |
| -sub-Gaussian | a mean-zero with for all ; this is not the regret | Section 13.2.2 |
The book writes with the variance as the second argument, as most statistics texts do, and states the standard deviation separately where it matters.
A.3 Gaussian processes #
| Symbol | Meaning | Introduced |
|---|---|---|
| the unknown objective (or latent utility) | Section 1.1 | |
| the domain of inputs | Section 1.1 | |
| the number of candidates, when the domain is a finite set | Section 12.4 | |
| Gaussian process with mean function and kernel | Section 7.3 | |
| kernel (covariance function) | Section 3.1.1, Section 7.1.2 | |
| , | lengthscale; one lengthscale per input (ARD) | Section 3.1.1, Section 9.2 |
| squared amplitude, for a stationary kernel | Section 7.2.1 | |
| observation noise variance | Section 2.6.4, Section 8.3 | |
| , | observed inputs and observed values | Section 8.1 |
| kernel matrix of the observed inputs, | Section 7.3, Section 8.1 | |
| covariances between and the observed inputs | Section 8.1 | |
| , | posterior mean (after observations) | Section 8.1, Section 11.2 |
| , | posterior variance of the latent value (after observations) | Section 8.1, Section 11.2 |
| weights in the posterior mean | Section 3.5.1, Section 8.2.1 | |
| number of features of a linear model; for random Fourier features, the number of frequencies drawn | Section 7.1, Section 10.4.3 | |
| , , | reproducing kernel Hilbert space (RKHS) of , its inner product, and its norm | Section 10.2 |
| bound on the RKHS norm; some papers bound , others | Section 10.2.4 | |
| , | integral operator of a kernel, and the weighting density it integrates against | Section 10.1.3 |
| , | eigenvalues and eigenfunctions of (Mercer's theorem) | Section 10.3 |
| , | spectral density of a stationary kernel; the same normalized to a probability density | Section 10.4 |
One collision is worth knowing about. , with no argument, is the noise variance, following Rasmussen and Williams (2006); its subscript stands for noise. , always written with its argument, is the posterior standard deviation after observations, following Srinivas et al. (2010). The argument tells them apart. Where both appear in one formula (Section 12.6), the noise variance is written . In Chapter 10, with an index is an eigenvalue and a weighting density; elsewhere is the inverse Mills ratio or a lapse rate, and the number of options in a query.
Sources cited in Section A.3 2
- Rasmussen and Williams (2006) Gaussian Processes for Machine Learning
- Srinivas et al. (2010) Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
A.4 Optimization and preferences #
| Symbol | Meaning | Introduced |
|---|---|---|
| , | a maximizer and the maximum | Section 6.4.2, Section 11.1 |
| the incumbent, the best value observed after evaluations | Section 12.2 | |
| an acquisition function, after observations | Section 11.2 | |
| , , | probability of improvement, expected improvement, upper confidence bound | Section 12.2, Section 12.3, Section 12.4 |
| , | exploration weight of UCB, written | Section 11.2.1, Section 12.4 |
| improvement margin in PI and EI (a choice of experiment in Section 6.4) | Section 12.2 | |
| , | instantaneous regret; cumulative regret over rounds | Section 13.1 |
| maximum information gain of a kernel after observations | Section 6.5.2 | |
| is preferred to | Section 16.3 | |
| or | latent utility of a person (the book uses when the GP machinery is shared) | Section 18.1 |
| (in a link) | noise scale of a person's evaluation of one option | Section 16.3 |
| scale of the logistic link, | Section 16.4 | |
| inverse Mills ratio , only in the chapters that define it | Section 17.3, Section 18.2 | |
| lapse rate, the probability that an answer is a random slip | Section 20.4.2 | |
| negative Hessian of the log-likelihood; for pairwise answers, a weighted graph Laplacian | Section 17.2, Section 18.2 | |
| expected utility of the best option | Section 19.4 | |
| number of options shown in one query (qEUBO) | Section 19.4 |
Regret bounds are stated with , which hides logarithmic factors, and every rate in the book is given with its assumptions and its regret unit (Section 29.4).
References
- (2006). Gaussian Processes for Machine Learning. MIT Press. Cited in §A.3
- (2010). Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. ICML 2010. Cited in §A.3