Bayesian Inference
Section 2.5 updated a belief about a coin flip by flip, on a grid of eleven possible biases, and ended with a warning: grids do not scale. One unknown on eleven values needs eleven numbers, but an unknown function needs a number for every input. The way out it named was to choose distributions whose updates have a closed form, so that a few numbers summarize the whole belief. Chapter 4 then supplied the main such distribution, the Gaussian, and showed that conditioning and multiplying Gaussians stays inside the family.
This chapter puts the two together. It treats Bayesian inference as a procedure: write down what you believe before the data, write down how the data arise, and let Bayes' rule produce what you should believe afterward. We work the procedure exactly for two models. The first is a coin with a continuous bias, where a Beta distribution plays the role the grid played before. The second is a straight line, and then a plane, with Gaussian uncertainty about their coefficients. Along the way we meet the questions a Bayesian optimizer asks of its model: what single value to report, what the next observation will be, which of two models the data favor, and what to do when no formula is available.
The line matters more than it seems. Its prior and posterior are Gaussians over two weights, and Chapter 7 turns the prior into a prior over whole functions by giving the line many more weights; Chapter 8 then does the same for the posterior. The Gaussian processes of the next part are this chapter's linear model with the number of features taken to infinity.
5.1 Prior, likelihood, posterior #
Section 2.5 named the four parts of Bayes' rule for an unknown that takes finitely many values. Most unknowns in this book are continuous: a coin's bias anywhere in , the slope of a line, the value of a function. The rule keeps its form, with densities in place of probabilities and an integral in place of the sum:
Here stands for everything unknown, possibly a vector, and for the data. The prior and the likelihood together make up the model. The posterior is what the model concludes, and the evidence is the probability the model gave to the data before seeing them.
5.1.1 A model is a program that generates data #
A reader who writes software may find it easiest to think of a model as a program that simulates data. The prior says how to draw the unknown, and the likelihood says how to draw data given the unknown. For a coin:
- Draw a bias from the prior.
- For each flip , draw heads with probability .
For a noisy straight line, with an input for each observation:
- Draw an intercept and a slope from the prior.
- For each input , compute the line's value there and add independent Gaussian noise.
Bayesian inference runs such a program backwards. Given the output, it asks which values of the hidden draws could have produced it, and how probable each one is. Writing the model as a simulator also gives a practical check that the prior says what we mean: run the program with draws from the prior and look at the fake data it produces. If the fake data look absurd, the prior is wrong, and no amount of later computation will fix it. Such simulations are called prior predictive checks, and Gabry et al. (2019) show how to make them a routine step of building a model.
The second step of each program assumes that the observations are independent once is known (Section 2.7). The likelihood of the whole data set then factorizes,
and implementations work with the sum of logs, which does not underflow. As Section 2.5.2 showed, independence also lets the data arrive one at a time: the posterior after the first observations serves as the prior for observation , and the order of arrival does not matter.
5.1.2 Closed forms and conjugacy #
The integral in the evidence is where continuous Bayesian inference gets hard. For most pairs of prior and likelihood it has no formula, and in more than a few dimensions it cannot be computed on a grid either. The proportional form of Bayes' rule, Equation (2.7), sidesteps the integral whenever we can recognize the shape of its right-hand side, likelihood times prior, as a known distribution, because then the normalizing constant is known too. Section 4.6 did this: the product of a Gaussian prior and a Gaussian likelihood had the shape of a Gaussian, so the posterior was a Gaussian with no integral computed.
A prior with this property has a name. A family of priors is conjugate to a likelihood when every posterior is again in the family. Updating a conjugate model means updating a few parameters, which is why the next two sections can show exact posteriors that change as you add data.
Everything a Bayesian model can tell us is computed from the posterior: a best guess, a range of plausible values, a prediction, a decision. Summaries throw information away, and the information they most often discard, how uncertain the model is, is the information an optimizer needs to choose its next query.
Sources cited in Section 5.1 1
- Gabry et al. (2019) Visualization in Bayesian Workflow
5.2 A conjugate example #
Return to the coin, now with a bias that can be any number in . The question is the one Thompson posed in 1933: given the successes and failures seen so far, what should we believe about an unknown success probability, and how likely is it that one such probability exceeds another (Thompson, 1933)? The same question is asked today of a button's click-through rate in an online experiment (Section 15.6) and of each option in a multi-armed bandit, the problem of choosing repeatedly among a few options ("arms") with unknown success rates (Section 13.2).
5.2.1 The Beta distribution #
We need a family of densities on flexible enough to express a prior and closed under the coin's update. The likelihood of heads and tails in a particular sequence of flips is . A prior of the same shape will multiply with it and keep its shape. That is the Beta distribution:
where is the constant that makes the area one, the Beta function. With the density is flat, the uniform prior. Larger and equal and give a bump centered at one half, and unequal ones tilt it. Its mean and, for , its mode are
and its variance is , which shrinks as grows.
- By Equation (5.2), heads and tails have likelihood .
- By the proportional form of Bayes' rule and Equation (5.3), . The constant does not involve and drops out.
- Adding exponents, this is .
- This has the shape of Equation (5.3) with parameters and . A density is determined by its shape, since the constant is whatever makes the area one, so .
- The evidence is the ratio of the constants: .
The update is counting. Add the heads to and the tails to .
5.2.2 Prior strength in units of data #
The update suggests a reading of the prior's parameters. A prior behaves as if we had already seen heads and tails, so we can describe it by two numbers that mean something: its mean , the bias we expect, and its strength , the number of flips it is worth. After real flips, the posterior mean is
The posterior mean is a weighted average of the prior mean and the observed fraction of heads, with weights in proportion to the prior's strength and the number of flips. With few flips the prior dominates; with many, the data do. Pulling an estimate toward a prior value in this way is called shrinkage, and Equation (5.4) says exactly how much to shrink: by the ratio of pretend data to real data.
The figure lets you set both halves and watch the posterior respond.
Start from the uniform prior. With the default and mean 0.5, the prior is , flat. The posterior then has exactly the shape of the likelihood, and the dotted and solid curves coincide: with a flat prior, Bayes' rule only normalizes the likelihood.
Make the prior strong. Raise to 50. Seven heads and three tails now barely move the posterior from 0.5: the prior is worth five times as many flips as the data. Then raise the heads to 70 and the tails to 30 and watch the data take over.
Make the prior wrong and confident. Set the prior mean to 0.2 with , and the data to 70 heads and 30 tails. The posterior settles between the two, closer to the data. A confident prior that conflicts with the data costs many observations to overcome, which is a reason to keep priors honest about what is not known.
Watch the interval shrink. Keep the ratio of heads to tails fixed and multiply both counts by four. The credible interval narrows by about half. The posterior standard deviation of a fraction falls like , so four times the data buys twice the precision.
The 95% credible interval shaded in the figure is the range that holds 95% of the posterior probability. It answers the question most people think a confidence interval answers: given these data, where does probably lie?
5.2.3 What the coin predicts #
The posterior also answers a question about the future: what is the probability that the next flip lands heads? Average the probability of heads, , over the posterior:
With the uniform prior this is , Laplace's rule of succession. After three heads in three flips it gives , not the certainty that the observed fraction would suggest. This averaging over the posterior is the general recipe for prediction, and Section 5.5 returns to it.
Two uses of this model recur later. Thompson sampling, which Section 12.5 develops for functions, began as a rule for this setting: draw a bias from each option's Beta posterior and try the option whose draw is largest (Thompson, 1933). And the model needs nothing but a yes or a no per observation, so it reappears whenever the data are binary answers. A 2024 system that recommends items through a dialogue, for example, keeps a Beta posterior for each item and updates it after each of the user's answers, by an amount that a language model assigns for how strongly the answer supports that item (Austin et al., 2024a).
Sources cited in Section 5.2 2
- Thompson (1933) On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples
- Austin et al. (2024a) Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation
5.3 Point estimates and their limits #
Often a single number is wanted: the bias of the coin, the slope of the line, the best setting of a learning rate. There are three standard ways to reduce a posterior to one value, and it is worth knowing what each one does before seeing what all of them lose.
The maximum likelihood estimate (MLE) ignores the prior and picks the value under which the data were most probable: . For the coin it is the fraction of heads, . For a line with Gaussian noise, maximizing the likelihood means minimizing the sum of squared residuals, so the MLE is the least-squares fit (Exercise 5.3).
The maximum a posteriori estimate (MAP) picks the peak of the posterior: . For the coin it is the posterior mode, . In log form, the MAP maximizes the log likelihood plus the log prior, and the log prior acts as a penalty. With a Gaussian prior on a line's weights the penalty is a multiple of the squared length of the weight vector, and the MAP is ridge regression, the least-squares fit with an penalty that many software engineers have met as regularization.
The posterior mean averages over the posterior. It is the estimate with the smallest expected squared error, and for the coin it is Equation (5.4). A reader who has used add-one smoothing to avoid zero counts has used the posterior mean under a uniform prior.
Read the MLE after three heads. It says the coin always lands heads, and a model built on it would bet anything against a tail. The posterior says something milder: the bias is probably high, but anywhere from about 0.4 to 1 is plausible.
Give the prior a little weight. Set the prior strength to 4 with mean 0.5, that is, . The MAP moves off the edge to 0.8 and the posterior mean to 0.71. The prior acts as two imaginary heads and two imaginary tails, a guard against overconfidence from little data.
Add data and watch the estimates converge. At 60 heads and 20 tails all three estimates lie within 0.02 of each other. With enough data the choice of estimate stops mattering, and so does the prior.
5.3.1 What a point estimate cannot do #
The deeper problem with any point estimate is not that it can be wrong. It is that it does not say how wrong it might be. A coin with 1 head in 2 flips and a coin with 500 heads in 1000 flips both have an MLE of 0.5. Anyone would trust the second estimate more, and the posterior records why: the first is under a uniform prior, nearly as wide as the prior itself, while the second has a standard deviation of about 0.016.
For optimization, this missing information is the whole game. Suppose two learning rates have been tried. The first was run three times, with a mean validation accuracy of 0.80; the second was run once and reached 0.78. A point estimate says the first is better and the second should be abandoned. The posterior says the second might well be better, because one noisy run leaves its accuracy uncertain, and that one more run there would teach more than a fourth run of the first. Choosing between exploiting what looks best and exploring what is uncertain is the central trade-off of Bayesian optimization (Section 11.3), and it can only be made by a model that keeps its uncertainty. Every acquisition function in Chapter 12 is a rule that reads both the posterior mean and the posterior spread.
5.4 Bayesian linear regression #
The coin had one unknown number. An objective function is an unknown relationship between inputs and outputs, and the simplest such relationship is a straight line. We observe noisy values at inputs and model them as
an intercept , a slope , and independent Gaussian noise with standard deviation . The question is what we should believe about the line, and therefore about its value at inputs we have not tried, after a few observations.
5.4.1 The model #
Collect the weights into , and collect what each weight multiplies, here the constant 1 and the input , into a vector of features , so that the line's value is . For observations, stack the feature vectors as the rows of the design matrix and the observations into . The model is:
The prior covariance says how large we expect the weights to be; in the figures it is , independent weights with standard deviation . Nothing in this section depends on the features being and . Any fixed list of features works, which is the door that Chapter 7 walks through.
Before any data, the prior over weights is already a prior over lines. Each draw of is one line, and the value at has variance , by Equation (4.9). The prior lines therefore fan out from the region around , where only the intercept varies, and spread quadratically away from it.
5.4.2 The posterior #
Both the prior and the likelihood are Gaussian in , so the posterior is too, and the recipe of Equation (4.6) finds it.
- By Bayes' rule, , where the constant is the log evidence, which does not involve .
- The log likelihood of Equation (5.6) is , and the log prior is .
- Expand the squared norm: . The first term does not involve .
- Collect the terms quadratic and linear in : .
- This is the form of Equation (4.6) with precision and linear coefficient . So the posterior is Gaussian with covariance and mean .
In one line:
This is the weight-space result of Gaussian process textbooks (Rasmussen and Williams, 2006, sec. 2.1), and it has the shape we met in Section 4.6.2. The posterior precision is the prior precision plus a term from the data, and the data term is a sum over observations, : each observation adds precision along its own feature vector and nowhere else. The posterior mean is the ridge regression fit with penalty (Exercise 5.3). As in Section 4.5, the posterior covariance does not depend on the observed values , only on where the observations were made.
The same posterior can be reached a second way. The pair is jointly Gaussian, since is a linear map of plus independent noise, so we could condition on with Equation (4.15). That route inverts an matrix, one row per observation, where Equation (5.7) inverts a matrix, one row per weight. The two answers agree, by an identity that turns the inverse of an matrix of this form into the inverse of a one (the Woodbury identity, Section B.2), and Chapter 7 follows the route because it survives when the number of weights becomes infinite.
5.4.3 Two views of the same posterior #
The figure shows the posterior in two spaces. On the left is data space, where each observation is a point and each weight vector is a line. On the right is weight space, where each weight vector is a point: the ellipses are the prior and the posterior over , drawn like the two-dimensional Gaussians of Figure 4.1. Each violet line on the left is one draw from the posterior, and the violet dot of the same draw sits on the right.
Clear the data. The posterior is the prior: a circle in weight space and, in data space, a fan of lines narrowest at .
Add a single point. In weight space the circle collapses into a long thin ellipse. One observation pins down the line's value at one input, a single combination of intercept and slope, and leaves the other combination free: in data space every drawn line passes near the point and pivots around it (Exercise 5.2).
Add a second point far from the first. The ellipse shrinks to a small blob and the lines nearly agree. Two well-separated points determine a line.
Now put the two points close together. The slope stays uncertain, the ellipse stays long, and the lines fan out on both sides of the pair. Where the observations sit matters as much as how many there are.
Raise the noise. Each observation adds less precision, the posterior relaxes toward the prior, and the lines miss the points more freely.
5.4.4 A plane, and more inputs #
Objectives in this book usually have more than one input, and a line becomes a plane. With two inputs, the features are and the weights are an intercept and two slopes. We write them , numbering from zero so that the slope carries the index of its input , and Equation (5.7) applies unchanged with a precision matrix. What changes is the geometry of what the data can determine, and the figure below makes it visible.
Look at the starting data. The three inputs lie on the diagonal . The drawn planes all pass near the three points, but they swing about the diagonal like a door on a hinge, and the map is light along the diagonal and dark in the two far corners. The data determine the slope along the diagonal and say nothing about the tilt across it; the readout names that direction.
Click the map near a far corner. One observation off the diagonal stops the swinging. The planes snap together and the map lightens everywhere.
Clear the map and evaluate random inputs. Three inputs in general position, not on one line, are enough to pin down a plane up to the noise.
The general rule follows from the structure of . Each observation adds precision only along its feature vector, so the data constrain only the directions in weight space that their inputs span. With inputs a linear model has weights, and it needs at least observations whose inputs do not all lie in a lower-dimensional flat before every direction is constrained. A six-parameter tuning problem fitted with a linear model needs seven well-spread evaluations before the model stops being as uncertain as its prior in some direction. A Gaussian process, which can bend, needs many more, and how that number grows with the dimension is a theme of Chapter 30.
Linear models in many dimensions are not only a teaching device. When the features are themselves learned by a neural network, Bayesian linear regression on those features has served as the surrogate for tuning image-recognition and image-captioning models at large scale, with a cost linear in the number of observations (Snoek et al., 2015). More strikingly, a 2026 study found that Bayesian linear regression, after a geometric transformation of the inputs, matched the leading high-dimensional Bayesian optimization methods on tasks with 60 to 6,000 dimensions (Doumont et al., 2026).
5.4.5 Computing it #
The cost of Equation (5.7) is to form for features and observations, and to factorize it. It grows only linearly with the number of observations, which is the property the two studies above exploit. A Gaussian process, by contrast, pays (Section 8.4).
import numpy as np
def blr_posterior(Phi, y, prior_var, noise_var):
"""Posterior mean and covariance of w for y = Phi w + noise, w ~ N(0, prior_var I)."""
M = Phi.shape[1]
A = Phi.T @ Phi / noise_var + np.eye(M) / prior_var # posterior precision
L = np.linalg.cholesky(A)
mean = np.linalg.solve(L.T, np.linalg.solve(L, Phi.T @ y / noise_var))
Linv = np.linalg.solve(L, np.eye(M))
cov = Linv.T @ Linv # A^{-1}
return mean, cov
x = np.array([-0.8, -0.3, 0.2, 0.7]); y = np.array([-0.5, 0.0, 0.35, 0.9])
Phi = np.column_stack([np.ones_like(x), x]) # features (1, x)
mean, cov = blr_posterior(Phi, y, prior_var=1.0, noise_var=0.09)
Sources cited in Section 5.4 3
- Rasmussen and Williams (2006) Gaussian Processes for Machine Learning
- Snoek et al. (2015) Scalable Bayesian Optimization Using Deep Neural Networks
- Doumont et al. (2026) We Still Don't Understand High-Dimensional Bayesian Optimization
5.5 The predictive distribution #
A posterior over weights is a means to an end. What an optimizer needs is a belief about the objective at an input it has not tried, and about the value it would observe there. Both come from the same recipe as the coin's next flip: average the prediction of each possible parameter value over the posterior,
This is the posterior predictive distribution. It accounts for two sources of uncertainty at once: we do not know the parameters exactly, and even if we did, a new observation would carry fresh noise.
For the linear model the integral needs no computing. The latent value is a linear map of the Gaussian posterior over , so by Equation (4.9) it is Gaussian, and the observation adds independent noise, which by Equation (4.18) adds its variance:
with . The first variance is the model's uncertainty about the line; the second adds the noise of one more measurement, the distinction drawn in Section 8.3.
Read the band's shape. It is narrowest at the center of the data and grows on both sides, a bow tie. The line is pinned near the data and free to pivot, so its uncertainty grows with the distance from the center.
Compare the two bands. Near the data the dashed band is much wider than the shaded one: there the uncertainty about a new observation is mostly measurement noise. Far away the two bands converge, because uncertainty about the line dominates.
Spread the points out. Add observations near both ends. The bow tie flattens, because well-separated inputs determine the slope.
A tempting shortcut is to predict with a single fitted line, plugging the point estimate into the likelihood. The resulting plug-in predictive has variance everywhere: it claims to know the line exactly, even at , far from any data. It is overconfident exactly where an optimizer most needs to know that it does not know.
The bow tie also shows what a linear model cannot represent. Between two clusters of data, where nothing has been observed, a line is still pinned by the clusters on either side, so its uncertainty there is small. For a straight line that is correct. For an unknown objective it is not: the function could do anything in the gap. A model for optimization needs uncertainty that grows wherever data are absent, not only far from their center, and Chapter 7 gets it by giving the linear model enough features to bend.
5.6 Model evidence #
Every model so far came with choices: the prior's mean and strength for the coin, the prior width and the noise for the line, the features themselves. Numbers of this kind, which set up the model and are not among the unknowns it reasons about, are the model's hyperparameters. (Machine learning uses the same word for the settings of a training procedure, as in Section 1.1. In both uses it means a setting one level above the quantities being learned.) How should the data inform such choices? Picking the model whose best fit has the highest likelihood does not work, because a more flexible model always fits at least as well, including fitting the noise.
Bayes' rule answers one level up. Treat the model itself as unknown, and its probability after the data is
The quantity that decides is the evidence of Equation (5.1), also called the marginal likelihood: the probability the model assigned to the data before seeing them, averaged over its prior. It does not reward the best fit the model could achieve, but the fit it predicted on average.
5.6.1 Occam's razor, automatically #
That average builds in a preference for simple models. A probability distribution over possible data sets must sum to one. A flexible model, able to explain many different data sets, spreads that probability thin, and gives each data set it could explain only a little. A rigid model concentrates its probability on fewer data sets. When the data fall among the rigid model's predictions, it wins; when they do not, the flexible model wins. This is Occam's razor as a consequence of probability rather than a rule of thumb (MacKay, 2003, ch. 28).
Compare two models of a coin. says it is fair, exactly. says the bias is unknown, with a uniform prior. For a particular sequence of flips with heads, gives probability , and by step 5 of the derivation in Section 5.2.1, gives .
With 5 heads in 10 flips, the evidence is for the fair coin and for the unknown bias: the data favor the fair coin by a factor of about 2.7. The flexible model spent probability on lopsided outcomes that did not happen. With 9 heads in 10 flips, the fair coin's evidence is unchanged while the flexible model's rises to , and the data favor an unknown bias by a factor of about 9.3.
For the linear model the evidence has a closed form. The observations are a linear map of the Gaussian weights plus independent Gaussian noise, so by Equation (4.9) and Equation (4.18) they are Gaussian, with , and the log evidence is the log density of that Gaussian at the observed :
The two terms pull in opposite directions. A wider prior makes larger, which improves the data fit term for data that need large weights but makes the determinant larger, the cost of spreading probability over more data sets. The same formula, with a kernel matrix in place of , is the marginal likelihood that Section 9.3 maximizes to fit a Gaussian process's hyperparameters.
Sweep the prior width. Move from 0.1 to 4 and watch the log evidence. It rises from about to a maximum near and falls again to about . Choosing at the peak is the practice called type II maximum likelihood, or empirical Bayes: maximize the evidence over the prior's settings rather than the likelihood over the weights.
Vary the noise too. A noise level far too small makes the data fit term punishing; one far too large makes every data set look plausible and none likely. The evidence is highest for the noise that explains the scatter of the points around a line.
The evidence has known weaknesses. It depends on the prior's width even where the posterior barely does: doubling an already vague prior hardly changes the fitted line but lowers the evidence. And maximizing it over many hyperparameters can overfit, that is, fit the accidents of a small data set, a failure Section 9.4 discusses for Gaussian processes. It remains the standard tool for setting the hyperparameters of the models in this book.
Sources cited in Section 5.6 1
- MacKay (2003) Information Theory, Inference, and Learning Algorithms
5.7 When there is no closed form #
Conjugate pairs are the exception. The Beta prior fitted the coin because the likelihood was a power of and ; the Gaussian prior fitted the line because the noise was Gaussian. Change either and the posterior leaves the family.
The case this book cares most about is a comparison. When a person says that option is better than option , a common model, developed in Section 16.3, sets the probability of that answer to , the standard normal distribution function from Equation (4.3) applied to the difference in utility. (Here is the noise of the difference; Section 16.3 writes it as times the noise of each option.) With a Gaussian prior on the difference , the posterior after one answer is proportional to
a Gaussian density multiplied by an S-shaped curve. The product is lopsided: it keeps the prior's upper tail and cuts away its lower one. It is not a Gaussian, and after many answers about many options, its normalizing integral runs over as many dimensions as there are options.
Four families of methods handle such posteriors, and Chapter 17 compares them. The Laplace approximation (Section 17.2) replaces the posterior by a Gaussian centered at its peak, with a covariance taken from the curvature there. Expectation propagation (Section 17.3) and variational inference (Section 17.4) choose a Gaussian by matching the posterior in other senses. Sampling methods (Section 17.5) draw from the posterior itself and replace every integral by an average over draws. The first three bring the problem back to Gaussians, so that everything in Chapter 4 applies again. The Gaussian process preference model of Chu and Ghahramani (2005), on which much of Part IV builds, used the Laplace approximation.
One tool is still missing. This chapter argued that an optimizer should evaluate where it will learn the most, without saying how learning is measured. Chapter 6 supplies the unit.
Sources cited in Section 5.7 1
- Chu and Ghahramani (2005) Preference learning with Gaussian processes
5.8 Exercises #
A coin has a prior, and you observe 8 heads in 10 flips. Find the posterior, its mean, its mode, and the maximum likelihood estimate. How many flips is the prior worth, and what is its weight in the posterior mean?
Solution
The posterior is . Its mean is and its mode is , while the MLE is . The prior is worth flips with mean 0.5, so by Equation (5.4) the posterior mean is : the prior carries a weight of .
Take the line model with and a single observation at input , so is the row . Show that the posterior precision leaves the prior variance unchanged in one direction of weight space. Which direction is it, and what does it mean for the lines in Figure 5.3?
Solution
Here . For any orthogonal to , , so is an eigenvector with eigenvalue 1, and the posterior variance in that direction equals the prior variance 1. Along itself the eigenvalue is , larger, so the variance shrinks. The constrained combination is , the line's value at . The unconstrained direction changes the slope by one unit and the intercept by , which rotates the line about the point . That is the pivoting of the drawn lines.
Show that with the posterior mean Equation (5.7) minimizes the ridge regression objective with . What does the posterior mean become as ?
Solution
Setting the gradient of the objective to zero gives , so . The posterior mean is ; multiplying inside and outside the inverse by turns it into , the same expression. Since the posterior is Gaussian, its mean is also its mode, so this is the MAP estimate too. As , and the posterior mean tends to the least-squares solution , the maximum likelihood estimate, provided is invertible.
Under a uniform prior, find the evidence for a particular sequence of 3 heads in 3 flips, and compare it with the fair-coin model of Example 5.1. Which model do the data favor, and by how much? How many heads in a row are needed before the unknown-bias model is favored by a factor of 10?
Solution
The unknown-bias model gives , and the fair coin gives , so the data favor the unknown bias by a factor of 2. For heads in a row the ratio is , which is for and for . Seven heads in a row are needed for a factor of 10. A run of heads is evidence against fairness, but less overwhelming than intuition suggests, because the flexible model had to spread its probability over every possible bias.
Further reading #
- Gelman et al. (2013), chapter 2, treats single-parameter models, starting with the binomial and its Beta prior, and the book as a whole is the standard reference for applied Bayesian modeling, including prior predictive checks.
- Bishop (2006), section 2.1, covers the Beta distribution and the coin, and chapter 3 covers Bayesian linear regression, its predictive distribution, and the evidence for model comparison in the notation used here.
- Rasmussen and Williams (2006), section 2.1, derives Bayesian linear regression in the weight-space view and then passes to Gaussian processes, the step Chapter 7 takes.
- MacKay (2003), chapter 28, explains Occam's razor through the evidence, with the pictures that made the argument famous.
- Thompson (1933) posed the question of comparing two unknown success probabilities from data, and with it the sampling rule that now carries his name.
References
- (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Cited in §5.2
- (2006). Pattern Recognition and Machine Learning. Springer.
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Cited in §5.7
- (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). Cited in §5.4
- (2019). Visualization in Bayesian Workflow. Journal of the Royal Statistical Society Series A: Statistics in Society. Cited in §5.1
- (2013). Bayesian Data Analysis. Chapman and Hall/CRC.
- (2003). Information Theory, Inference, and Learning Algorithms. Cambridge University Press. Cited in §5.6
- (2006). Gaussian Processes for Machine Learning. MIT Press. Cited in §5.4
- (2015). Scalable Bayesian Optimization Using Deep Neural Networks. Proceedings of the 32nd International Conference on Machine Learning (ICML 2015). Cited in §5.4
- (1933). On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika. Cited in §5.2