When the Posterior Is Not Gaussian
Gaussian process regression owes its tidy formulas to one fact: a Gaussian prior combined with Gaussian observation noise gives a Gaussian posterior. The comparison models of Chapter 16 break that fact. Their likelihood is a probit or logistic curve, not a Gaussian, and the posterior over the utility is no longer Gaussian. Everything downstream, the predictive mean and band, the probability of the next answer, and the acquisition functions of Chapter 19, needs that posterior.
This chapter works almost entirely with the smallest case that shows the difficulty: one utility difference, a Gaussian prior, and a few comparisons. Its exact posterior can be computed on a grid and drawn, so each approximation can be checked against the truth. We then move to two and more latent values, where the exact answer turns out to be skewed only in the directions that comparisons touch, and end with the evidence on how much the choice of approximation matters in practice.
17.1 Non-Gaussian likelihoods #
Recall why regression was easy. With a prior and observations with Gaussian noise, Bayes' rule multiplies two Gaussian densities in . The product of Gaussian densities is again a Gaussian density up to a constant (Section 4.6), so the posterior has a closed form and so does its normalizing constant, the marginal likelihood. Prior and likelihood form a conjugate pair.
A comparison does not. Write for the utilities at the inputs that have been compared, and suppose answers have been recorded, the -th saying that input was preferred to input . With the probit model of Equation (16.2), Bayes' rule gives
The prior is Gaussian, but each likelihood factor is an S-shaped curve in one direction of , the direction of the difference . The product is not Gaussian, and is the probability that a Gaussian vector lands in a region bounded by soft walls, an -dimensional integral with no closed form in general (Section 5.7). Gaussian process classification, in which each input carries a yes-or-no label instead of a number, has the same structure, with one factor per labeled input (Rasmussen and Williams, 2006, ch. 3), so the methods of this chapter were developed and tested there first.
17.1.1 The smallest case #
Strip the problem down to one number. Let be the utility difference between two options, as in Section 16.6, with a Gaussian prior . Here is a variance; Equation (16.6) described the same kind of belief by its standard deviation , so . Suppose a person says once that is better. Write for the noise on the difference. The posterior is
The factor 2 is : by Equation (16.6) with a belief centered at zero, the prior probability of the answer is . A Gaussian density multiplied by a normal distribution function is called a skew-normal density, a family introduced by Azzalini (1985). It keeps the upper tail of the prior and cuts away the lower one, softly when the noise is large and sharply when it is small. Its mean has a closed form.
- For and a differentiable , integration by parts gives , because the derivative of the density is . This is Stein's lemma.
- Take , so . Then .
- is the density seen as a function of , so its prior expectation is the density at 0 of with independent (Section 4.6), namely .
- Divide by : the posterior mean is .
With , the mean grows from about 0.46 at to as the noise vanishes, and the standard deviation falls from about 0.89 to 0.60 (Exercise 17.1). The mode, the peak of the density, behaves differently: it solves , and as shrinks it slides toward zero, to 0.05 at . In the noise-free limit the posterior is the prior's upper half, a half-normal, whose peak sits at the cut while its mass lies well to the right of it.
Some things to try:
- Shrink the noise. Drag toward 0.01. The likelihood becomes a step, the posterior becomes the half-normal, and the mode and the mean separate: the mode goes to the cut, the mean stays near 0.80.
- Raise the noise. At the likelihood is a gentle slope, the posterior is a slightly shifted bell, and mode and mean nearly coincide.
- Let both win. With three wins for and three for the posterior is pinned from both sides and is close to a Gaussian centered at zero. Skew comes from one-sided evidence.
- Let keep winning. Each further win pushes the mass up, but with small noise the left edge stays sharp: after a run of identical answers the evidence says " is better" more firmly, not "by how much".
This lopsided shape is what every method below has to summarize, usually with a single Gaussian.
Sources cited in Section 17.1 2
- Rasmussen and Williams (2006) Gaussian Processes for Machine Learning
- Azzalini (1985) A Class of Distributions Which Includes the Normal Ones
17.2 The Laplace approximation #
The simplest summary puts a Gaussian at the posterior's peak and gives it the peak's curvature. A bell shape is determined by where it is highest and how fast it falls off from there; if the posterior is close to a bell, those two facts recover it. This is the Laplace approximation, named after Laplace's method for approximating integrals, and Tierney and Kadane (1986) showed how far it goes for Bayesian computation: approximate posterior moments and marginal densities need only a maximization and the curvature at the maximum.
Write for the log of the unnormalized posterior, and for its maximizer, the mode or maximum a posteriori estimate.
- Write for the gradient of , the vector of its first derivatives with respect to the entries of , and for its Hessian, the matrix of its second derivatives. Expand to second order around (Taylor's theorem): , with .
- At a maximum the gradient vanishes, , so the linear term drops.
- The Hessian of the log prior is . Write at for the curvature of the log likelihood; then .
- Exponentiate: , which is a Gaussian density up to a constant.
- Hence .
Two things are needed: the mode and the curvature there. For the probit likelihood the log likelihood is concave, because is concave, and the log prior is a concave quadratic, so has a single maximum and Newton's method finds it quickly. Newton's method replaces the function by its quadratic approximation at the current point and jumps to that quadratic's maximum.
- The gradient is and the Hessian is , with now evaluated at the current .
- The maximizer of the local quadratic is .
- Write and collect terms: .
- Avoid inverting : since , we have , so .
Starting from and repeating Equation (17.4), halving the step whenever fails to increase, converges in a handful of iterations. For classification, where is diagonal, Rasmussen and Williams (2006) give a numerically stable version as their Algorithm 3.1. For comparisons is not diagonal; Section 18.2 works out its structure.
The same expansion approximates the normalizing constant , the marginal likelihood used to fit hyperparameters (Section 9.3). Integrating the Gaussian of step 4 gives
the Laplace evidence that Chu and Ghahramani (2005) used for preferences (their equation 12) and that BoTorch maximizes to fit its preference model (Section 18.6).
17.2.1 Where it goes wrong #
The figure below adds the Laplace Gaussian to the exact posterior.
At moderate noise the blue curve sits close to the exact one. As the noise shrinks it fails in a specific way. At the exact posterior has mean 0.80, standard deviation 0.60, and gives a probability of 0.004 of being better. The Laplace Gaussian is centered at the mode, 0.05, with standard deviation 0.27, and gives a probability of 0.43: after a single nearly noise-free answer it is almost as unsure of the order as before (Exercise 17.2). The mode sits at the edge of the mass, and a Gaussian centered there must spill over the edge.
The general lesson was established for classification. Kuss and Rasmussen (2005) compared Laplace's method and expectation propagation with long sampling runs on binary Gaussian process classifiers and found that Laplace's method "systematically underestimates the mean", so that the approximate posterior over latent functions has "too small amplitude" and its predictive probabilities are over-conservative, although the sign of the latent function is mostly right; they concluded that it is "so inaccurate that we advise against its use, especially when predictive probabilities are to be taken seriously." The same pattern appears for comparisons: with nearly noise-free duels, the mode can be very far from the mean (Takeno et al., 2023).
Laplace's method survives because it is fast and simple. In robot reward learning from pairwise preferences, Bıyık et al. (2020) described expectation propagation as "more accurate than Laplace approximation" but "slower in practice", and chose Laplace for its computational efficiency. It is also the default in BoTorch (Section 18.6).
Sources cited in Section 17.2 6
- Tierney and Kadane (1986) Accurate Approximations for Posterior Moments and Marginal Densities
- Rasmussen and Williams (2006) Gaussian Processes for Machine Learning
- Chu and Ghahramani (2005) Preference learning with Gaussian processes
- Kuss and Rasmussen (2005) Assessing Approximate Inference for Binary Gaussian Process Classification
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Bıyık et al. (2020) Active Preference-Based Gaussian Process Regression for Reward Learning
17.3 Expectation propagation #
The Laplace approximation looks at one point of the posterior. A better Gaussian would match the posterior's mean and variance, its mass rather than its peak. Computing those moments for the full posterior is as hard as the original problem, but computing them for a posterior with only one non-Gaussian factor is easy. Expectation propagation (EP) builds the approximation from such one-factor problems (Minka, 2001).
EP replaces each likelihood factor , called a site, by an unnormalized Gaussian in the same direction, so that the approximate posterior is Gaussian. It then refines one site at a time.
Input: prior , sites .
- Initialize every site approximation to a constant, so that is the prior.
- Pick a site and remove its approximation, forming the cavity , a Gaussian.
- Multiply in the exact factor, forming the tilted distribution .
- Compute the mean and covariance of .
- Choose the new so that has exactly those moments, and update .
- Repeat steps 2 to 5 over all sites, in sweeps, until the site approximations stop changing.
The tilted distribution in step 3 has one non-Gaussian factor, which acts along one direction, so its moments reduce to a one-dimensional problem. For a probit site the answer is in closed form. Let the cavity give the difference the Gaussian , with a variance like , and let the site be with if won and if won.
- The normalizer is with , by the argument of Equation (16.6).
- Differentiating under the integral, , so , and the tilted mean is .
- Differentiating once more gives the tilted variance .
- With , the chain rule gives and , using .
- Hence the tilted mean is and the tilted variance is .
These are the same quantities, with , as in the expectation propagation algorithm for probit classification (Rasmussen and Williams, 2006, ch. 3). Each site update is a rank-one change to the covariance, which costs for latent values, so a sweep over sites costs , comparable to a Newton iteration when is of the order of .
With a single comparison there is a single site, the tilted distribution is the exact posterior, and EP returns its exact mean and variance: at , mean 0.80 and standard deviation 0.60. That is far better than Laplace's 0.12 and 0.32. Matching moments has a cost of its own, though. A Gaussian with the right mean and variance still puts mass where the exact posterior has none: EP gives a probability of 0.093 of being better, against an exact value of 0.013. With five wins at , several sites act on the same direction and EP is no longer exact (mean 0.93 and standard deviation 0.45, against 0.90 and 0.58).
That small misplacement matters for comparisons specifically. Takeno et al. (2023) took a long run of a sampling method as ground truth (10,000 samples from the Gibbs sampler of Section 17.5, after discarding the first 1,000 and keeping every tenth), and found that EP estimates means and credible intervals very accurately but, for a duel already observed with beating , overestimates the probability that . On the Ackley function, a standard test function, it underestimated duel probabilities near 0 or 1, and its estimates did not keep the true ordering between pairs, which can change which pair an acquisition function picks.
For classification the verdict on EP is favorable. Kuss and Rasmussen found its predictive probabilities and marginal likelihood estimates very close to those of long sampling runs (Kuss and Rasmussen, 2005), and a broader comparison of approximations for binary Gaussian process classification concluded that "the Expectation Propagation algorithm is almost always the method of choice unless the computational budget is very tight" (Nickisch and Rasmussen, 2008). Unlike Newton's method on a concave function, EP is not guaranteed to improve an objective at every step, so implementations damp the site updates when they oscillate.
Sources cited in Section 17.3 5
- Minka (2001) Expectation Propagation for Approximate Bayesian Inference
- Rasmussen and Williams (2006) Gaussian Processes for Machine Learning
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Kuss and Rasmussen (2005) Assessing Approximate Inference for Binary Gaussian Process Classification
- Nickisch and Rasmussen (2008) Approximations for Binary Gaussian Process Classification
17.4 Variational inference #
A third route turns approximation into optimization. Pick a family of simple distributions, Gaussians here, and find the member closest to the posterior, measuring closeness by the Kullback-Leibler divergence (Section 6.2). The divergence to the posterior cannot be computed directly, because it contains the unknown , but it can be minimized anyway.
- By Bayes' rule, .
- Take the expectation under of : .
- The first expectation is to the prior. Rearranging, .
- The last term is nonnegative, so , and since does not depend on , maximizing minimizes the divergence to the posterior.
is the evidence lower bound (ELBO). For a Gaussian and a probit likelihood, every term is cheap: the divergence between two Gaussians has a closed form, and the expected log likelihood is a sum of one-dimensional integrals, one per comparison, each over the Gaussian that assigns to a difference, so standard gradient-based optimizers can maximize it.
The direction of the divergence decides the character of the fit. averages under , so it is enormous wherever puts mass and has almost none. The optimal therefore stays inside the posterior's support and tends to be too narrow. Expectation propagation's local moment update instead chooses the Gaussian closest to the tilted distribution in the other direction, , which punishes for missing mass and tends to make it too wide.
At the variational Gaussian has mean 0.88 and standard deviation 0.30, against the exact 0.80 and 0.60, and gives a probability of 0.002 of being better, close to the exact 0.004. It respects the hard edge that EP crosses, but it halves the uncertainty. Each method errs where its criterion is blind.
Variational inference is the approximation that scales. With inducing points, a small set of pseudo-inputs that summarize the function (Titsias, 2009), and stochastic optimization, it handles thousands of comparisons and any likelihood whose expected log can be estimated. That is why it underlies much of the recent preference modeling: crowd preference learning with thousands of users and items (Simpson and Gurevych, 2020), top- rankings (Nguyen et al., 2021), choice functions (Benavoli et al., 2023), response times (Shvartsman et al., 2024), mixed likelihoods that combine comparisons with confidence ratings, in a 2025 preprint (Wu et al., 2025a), and the experiments of qEUBO (Astudillo et al., 2023). As of September 2026 we found no systematic assessment of its error for preference likelihoods comparable to those for Laplace and EP.
Sources cited in Section 17.4 7
- Titsias (2009) Variational Learning of Inducing Variables in Sparse Gaussian Processes
- Simpson and Gurevych (2020) Scalable Bayesian preference learning for crowds
- Nguyen et al. (2021) Top-$k$ Ranking Bayesian Optimization
- Benavoli et al. (2023) Learning Choice Functions with Gaussian Processes
- Shvartsman et al. (2024) Response Time Improves Gaussian Process Models for Perception and Preferences
- Wu et al. (2025a) Mixed Likelihood Variational Gaussian Processes
- Astudillo et al. (2023) qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization
17.5 Sampling #
The three methods so far replace the posterior by a Gaussian. Markov chain Monte Carlo (MCMC) methods instead draw a sequence of samples whose distribution converges to the posterior itself. Any quantity of interest, the mean, a credible interval, the probability that one option beats another, is then estimated by averaging over the samples. The answer becomes exact as the number of samples grows, at the price of computation and of having to judge when a chain has run long enough.
For a Gaussian prior times a likelihood, elliptical slice sampling is the natural choice. Murray et al. (2010) designed it for "models with multivariate Gaussian priors": it has "simple, generic code", "no free parameters" to tune, and "works well for a variety of Gaussian process based models". Its idea is to move along an ellipse that passes through the current state and a fresh draw from the prior, so every proposal is already plausible under the prior and only the likelihood needs checking.
Input: current state , prior , log likelihood .
- Draw , which defines the ellipse .
- Draw uniform on and set the threshold .
- Draw uniform on and set the bracket .
- If , accept and stop.
- Otherwise shrink the bracket toward zero, replacing the end on the same side of zero as , draw a new uniformly in it, and return to step 4.
The shrinking bracket always contains , the current state, so the loop terminates. Comparisons add a refinement. Because the probit factor is the probability that a Gaussian noise variable falls below the utility difference, the exact posterior is the marginal of a Gaussian restricted to a region cut out by linear constraints. Benavoli et al. (2021c) sample it with LinESS, a rejection-free elliptical slice sampler for linearly truncated Gaussians, at a cost of time and memory in the number of comparisons. Takeno et al. (2023) used Gibbs sampling, which redraws one variable at a time from its distribution given all the others, for the same distribution and found it faster: about 0.54 against 1.61 seconds in their Table 1, though LinESS should win when the number of truncations far exceeds the dimension.
With 200 samples the histogram is ragged and, with the default seed, the mean is off by almost 0.2: successive samples of a Markov chain are correlated, so 200 of them carry less information than 200 independent draws. With 1,000 the error is under 0.05, and with 5,000 the estimates agree with the exact column to about two digits. Sampling is the only method in this chapter whose error can be driven to zero by spending more computation. That is why studies of the other methods use it as the reference.
Sources cited in Section 17.5 3
- Murray et al. (2010) Elliptical Slice Sampling
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
17.6 An exact answer: the skew Gaussian process #
The one-comparison posterior Equation (17.2) has a name and a closed form. Does the general posterior Equation (17.1) have one too? It does, and the reason is the same random-utility trick used throughout Chapter 16: a probit factor is the probability that a hidden Gaussian variable is positive.
- For each comparison , write , so that , and introduce independent . Then .
- Stack the into an matrix and the into . By independence, the likelihood is , all inequalities holding at once.
- By Bayes' rule, is the distribution of given the event , where is jointly Gaussian.
- The normalizer is the probability of that event. The vector is Gaussian with mean zero and covariance , so is the probability that an -dimensional Gaussian vector is positive in every coordinate, an orthant probability.
A Gaussian vector conditioned on a linear transformation of itself, plus independent Gaussian noise, being positive is a unified skew-normal distribution. Durante (2019) showed that the posterior of parametric probit regression belongs to this family for any Gaussian prior, which makes the family conjugate to the probit likelihood. For functions, Benavoli et al. (2021c) proved that "the true posterior distribution of the preference function is a Skew Gaussian Process (SkewGP), with highly skewed pairwise marginals", and argued from it that Laplace's method "usually provides a very poor approximation". Theorem 29.1 states the theorem with its parameters.
The derivation also says where the posterior is skewed. Every factor depends on only through the differences . In any direction orthogonal to all the , the likelihood is constant and the posterior is the Gaussian prior conditioned on the differences. The skew lives in at most directions, and never in the direction that shifts every utility by the same amount, because no difference changes along it. The figure below shows the smallest instance: two utilities and comparisons between them.
Some things to try:
- Look along the diagonal. The exact contours are cut off by the line and extend freely along it. The skew is entirely across the diagonal, in the difference; along the diagonal, in the sum, the posterior is the Gaussian prior.
- Look at the top strip. The marginal of mixes the skewed difference with the Gaussian sum. At and the difference has skewness about 0.98, while alone has about 0.04. A single utility can look Gaussian even when the posterior is far from it. Kuss and Rasmussen (2005) made the same observation for classification: marginals of a high-dimensional truncated Gaussian "can be relatively similar to a Gaussian".
- Raise the correlation. Inputs close together in a kernel's eyes have correlated prior utilities, so their difference has a small prior variance, . The same noise is then large relative to that spread, the cut is softer, and the skewness of the difference falls, from about 0.98 at to 0.90 at . Two options the model already considers similar are also the ones a comparison moves least.
- Raise the noise. At the cut softens and the contours are nearly elliptical; the skewness of the difference drops to about 0.06.
The lesson carries to many dimensions. In a session of 30 comparisons among 60 distinct designs, each a point in a six-dimensional parameter space, the posterior lives on 60 latent utilities. It can be skewed only within the span of the 30 comparison directions, and the input dimension enters only through the kernel, which sets how correlated the utilities are. Laplace and EP both treat the remaining directions exactly. What they get wrong is the shape across the comparisons, and that is exactly what the probability of the next answer depends on (inference).
The exact posterior has a cost. Every prediction requires samples from a truncated multivariate Gaussian, and the marginal likelihood requires high-dimensional normal orthant probabilities (Benavoli et al., 2021c). Those are the reasons the approximations remain in use.
Sources cited in Section 17.6 3
- Durante (2019) Conjugate Bayes for probit regression via unified skew-normal distributions
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
- Kuss and Rasmussen (2005) Assessing Approximate Inference for Binary Gaussian Process Classification
17.7 How much the approximation matters #
The figures show what each approximation gets wrong in the smallest case. Does it change what a preferential optimizer does with real answers? Table 17.1 summarizes the methods before the evidence.
| Method | What it matches | Cost per fit | Typical failure | Used by |
|---|---|---|---|---|
| Laplace | the mode and the curvature there | a few Newton steps, each | centered at the edge of the mass when answers are nearly noise-free; mean and spread too small | Chu and Ghahramani; BoTorch PairwiseGP; Bıyık et al. 2020 |
| Expectation propagation | site by site, the moments of one factor at a time | sweeps of rank-one updates, per sweep | mass on the impossible side; wrong order of duel probabilities near 0 or 1 | Siivola et al. 2021; optuna-dashboard |
| Variational (Gaussian) | the Gaussian with the highest evidence lower bound | an optimization; scales with inducing points | too narrow | crowdGPPL; top- ranking; qEUBO experiments |
| Sampling | the posterior itself, in the limit | many samples; truncated Gaussians | correlated samples, so long runs; Monte Carlo error | Benavoli et al. 2021; Takeno et al. 2023 |
The strongest evidence comes from Takeno et al. (2023), with Gibbs sampling as ground truth, an RBF kernel, noise variance , and uniformly random duels. Laplace was inaccurate "since the mode can be very far away from the mean, particularly when" the noise variance is small, and they concluded that "although LA is very fast, LA-based preferential BO will fail". EP was very accurate for means and credible intervals but distorted duel probabilities, as described in Section 17.3. They also disagreed with the earlier test of Benavoli et al. (2021c), in which all inputs in one interval lose and all inputs in another win; they called such biased training duels "unrealistic" and used random duels instead. Even so, they fitted hyperparameters with the Laplace evidence, which they considered accurate enough at small noise.
These results come from nearly noise-free comparisons, the regime in which Figure 17.2 shows Laplace at its worst. Human comparisons are noisy, and as the figures show, noise softens the cut and shrinks the skew. How large the errors of Laplace and EP are at realistic human noise levels is a question no paper we found answers, and we found no independent replication of Takeno et al.'s comparison on real human comparisons (inference; both as of September 2026). Section 27.4 reports this evidence in full, with the libraries' choices.
In practice, three questions decide the choice. If the posterior feeds an acquisition function through probabilities of pairwise outcomes, as EUBO (Section 19.4) does, the misplaced mass of the Gaussian approximations matters most, and sampling or at least EP deserves the extra cost (inference). If only the posterior mean is needed, to recommend a final design, every method that gets the order of the utilities right is adequate. And if the session is long or the model has many users, variational inference is the method that scales.
Sources cited in Section 17.7 2
- Takeno et al. (2023) Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes
- Benavoli et al. (2021c) Preferential Bayesian optimisation with skew gaussian processes
17.8 Exercises #
Use Stein's lemma twice to show that the one-comparison posterior Equation (17.2) has second moment , the same as the prior, and hence variance . Check the values 0.89 at and 0.60 as for .
Solution
Apply Stein's lemma with : under the prior. The first term is . In the second, is proportional to a Gaussian density in with mean zero, so the expectation of times it vanishes. Dividing by gives . The variance is . With and , and the variance is , a standard deviation of about 0.89. As the variance tends to , a standard deviation of about 0.60. The comparison moves the mean but leaves the second moment alone: it reshapes the prior's mass without making it smaller.
For the one-comparison posterior with , show that as the Laplace approximation assigns probability tending to to the event , even though the exact probability tends to zero. You may use that the inverse Mills ratio satisfies for large .
Solution
The mode solves . Write ; then , so grows without bound as , but only like , and . The curvature is , using , so the Laplace standard deviation is about . The Laplace probability of is , and because grows only logarithmically. So the probability tends to . The exact posterior is the half-normal, which has no mass below zero. A noise-free answer that settles the order completely is reported by Laplace as no information about the order at all.
Let the target be the half-normal for and zero otherwise. (a) Explain why is infinite for every Gaussian . (b) Explain why the Gaussian minimizing has the mean and variance of . (c) Relate both answers to what Figure 17.4 and Figure 17.3 show at very small noise.
Solution
(a) , and every Gaussian puts positive mass on , where , so the divergence is infinite. At small but positive noise, is tiny but not zero there, and the best Gaussian under this divergence squeezes almost all its mass into , becoming narrow. (b) ; only the second term depends on , and for a Gaussian , which is maximized by and . (c) The variational fit in Figure 17.4 keeps out of and is too narrow; EP's moment matching in Figure 17.3 has the right mean and variance but spills over the edge.
Further reading #
- Rasmussen and Williams (2006), chapter 3, derives the Laplace approximation and expectation propagation for Gaussian process classification, with stable algorithms that carry over to comparisons.
- Kuss and Rasmussen (2005) and Nickisch and Rasmussen (2008) are the careful comparisons of approximations for classification, explaining why Laplace fails and EP does well.
- Minka (2001) introduces expectation propagation; Murray et al. (2010) introduces elliptical slice sampling; Tierney and Kadane (1986) is the classic on Laplace's method for posterior moments.
- Durante (2019) shows conjugacy of the probit model with unified skew-normal distributions; Benavoli et al. (2021c) extends it to the preference posterior and samples it exactly.
- Takeno et al. (2023) measure how wrong Laplace and EP are on duels and propose a cheaper exact-sampling alternative.
References
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Cited in §17.4
- (1985). A Class of Distributions Which Includes the Normal Ones. Scandinavian Journal of Statistics. Cited in §17.1
- (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. Cited in §17.2
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Cited in §17.2
- (2019). Conjugate Bayes for probit regression via unified skew-normal distributions. Biometrika. Cited in §17.6
- (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. Cited in §17.2 §17.3 §17.6
- (2001). Expectation Propagation for Approximate Bayesian Inference. Proceedings of the 17th Conference on Uncertainty in Artificial Intelligence (UAI 2001). Cited in §17.3
- (2010). Elliptical Slice Sampling. Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS 2010). Cited in §17.5
- (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Cited in §17.4
- (2008). Approximations for Binary Gaussian Process Classification. Journal of Machine Learning Research. Cited in §17.3
- (2006). Gaussian Processes for Machine Learning. MIT Press. Cited in §17.1 §17.2 §17.3
- (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Cited in §17.4
- (2020). Scalable Bayesian preference learning for crowds. Machine Learning. Cited in §17.4
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Cited in §17.2 §17.3 §17.5 §17.7
- (1986). Accurate Approximations for Posterior Moments and Marginal Densities. Journal of the American Statistical Association. Cited in §17.2
- (2009). Variational Learning of Inducing Variables in Sparse Gaussian Processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics (AISTATS 2009). Cited in §17.4
- (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. preprint Cited in §17.4