Learning from Comparisons
When the objective lives in a person's head, the most reliable measurement is often a comparison: this one or that one. This part rebuilds Bayesian optimization around that measurement. It starts with why comparisons work and the century-old models that turn a choice into evidence about a hidden utility, then extends the Gaussian process to learn from them, which makes the posterior non-Gaussian and calls for approximate inference.
With a preference model in hand, the part builds preferential Bayesian optimization itself: how to choose the next pair, why the expected utility of the best option is a principled answer, and what a person in the loop changes. It closes with the other questions a system can ask besides "which of two", and with the theory of dueling bandits behind all of it. In the middle of the part, you become the person being optimized.
The part assumes Part II and Part III; Chapter 16 can be read on its own.
Chapters in this part
- 16 Why Ask for Comparisons
Why a comparison is often a better measurement of a person than a rating, and the models that turn one into evidence about a hidden utility: psychophysics, Thurstone's comparative judgment, Bradley-Terry-Luce, random utility, how much one answer can carry, and the assumptions to watch.
- 17 When the Posterior Is Not Gaussian
Comparisons make the posterior non-Gaussian. Using one utility difference whose exact posterior can be drawn, the chapter derives and compares the Laplace approximation, expectation propagation, variational inference, and sampling, shows that the exact answer is a skew-normal (a skew Gaussian process in general) whose skew lives only along compared directions, and reports how much the choice matters.
- 18 Gaussian Process Preference Learning
Chu and Ghahramani's model: a Gaussian process utility observed only through noisy comparisons, fitted by Newton's method with the Laplace approximation. The chapter derives the fit step by step, predicts new comparisons, shows the model in one, two, and more dimensions, explains what comparisons cannot identify and what the comparison graph does to the posterior, and opens BoTorch's PairwiseGP.
- 19 Preferential Bayesian Optimization
Finding the best option from duels alone: the dueling formulation, how to pick the next pair, the decision-theoretic acquisition EUBO and qEUBO, its form for queries of several options, a complete loop with a person or a simulated one, and the failure modes reported in 2026.
- 20 Designing the Question
A pair is not the only question a system can ask. Choices among several and rankings, a slider that searches along a line, galleries and projections, answers that say 'about the same', 'not sure', or 'it crashed', many people at once, and why the interface belongs to the model.
- 21 Dueling Bandits and the Theory of Comparisons
The bandit view of learning from duels: what 'the best option' means when preferences are not transitive, the classic algorithms and their guarantees, the kernelized bounds of 2021 to 2026 with their assumptions and units, and the lower bound nobody has proved.
References for Part IV
141 works cited across this part's chapters.
- (2021). Instance-Wise Minimax-Optimal Algorithms for Logistic Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 21
- (2023). Identifying Nontransitive Preferences. University of Zurich. working paper Ch. 16
- (2018). Monotone Stochastic Choice Models: The Case of Risk and Time Preferences. Journal of Political Economy. Ch. 16
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Ch. 17 Ch. 19 Ch. 20
- (1985). A Class of Distributions Which Includes the Normal Ones. Scandinavian Journal of Statistics. Ch. 17
- (2025). A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures. Nature Human Behaviour. doi:10.1038/s41562-024-02085-2. Ch. 16
- (2020). BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). Ch. 19
- (2018). Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences. Nature Communications. Ch. 16
- (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. Ch. 21
- (2017). Noisy preferences in risky choice: A cautionary note. Psychological Review. Ch. 16
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Ch. 20
- (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. Ch. 17 Ch. 18
- (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika. Ch. 16 Ch. 20
- (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Ch. 18 Ch. 19 Ch. 20
- (2022). Can Market Participants Report Their Preferences Accurately (Enough)? Management Science. Ch. 16
- (2018). Predictably intransitive preferences. Judgment and Decision Making. Ch. 16
- (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. Ch. 19
- (2017). Sequential effects in preference decision: Prior preference assimilates current preference. PLOS ONE. Ch. 16
- (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. Ch. 18 Ch. 21
- (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Ch. 20
- (2021). Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance. Computer Graphics Forum. doi:10.1111/cgf.14188. Ch. 20
- (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. Ch. 21
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Ch. 16 Ch. 17 Ch. 18 Ch. 19 Ch. 21
- (1961). The Greatest of a Finite Set of Random Variables. Operations Research. Ch. 19
- (2021). Assessing Top- Preferences. ACM Transactions on Information Systems. Ch. 16
- (2006). Elements of Information Theory. Wiley. Ch. 20
- (2025). Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback. International Conference on Machine Learning. Ch. 21
- (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. preprint Ch. 20
- (2015). Contextual Dueling Bandits. Conference on Learning Theory. Ch. 21
- (2019). Conjugate Bayes for probit regression via unified skew-normal distributions. Biometrika. Ch. 17
- (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Ch. 16
- (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Ch. 20
- (2020). Improved Optimistic Algorithms for Logistic Bandits. International Conference on Machine Learning. Ch. 21
- (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. preprint Ch. 19
- (1860). Elemente der Psychophysik. Breitkopf und Härtel. Ch. 16
- (1973). Algebraic Connectivity of Graphs. Czechoslovak Mathematical Journal. Ch. 18
- (2014). The Limits of Attraction. Journal of Marketing Research. Ch. 20
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 19 Ch. 20 Ch. 21
- (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. Ch. 16
- (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. Ch. 18
- (1910). The Central Tendency of Judgment. The Journal of Philosophy, Psychology and Scientific Methods. Ch. 16
- (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Ch. 16 Ch. 18
- (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. Ch. 20
- (1982). Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis. Journal of Consumer Research. Ch. 20
- (2015). Sparse Dueling Bandits. Proceedings of the 18th International Conference on Artificial Intelligence and Statistics. Ch. 21
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Ch. 21
- (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. Ch. 21
- (2015). Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem. Conference on Learning Theory. Ch. 21
- (2016). Copeland Dueling Bandit Problem: Regret Lower Bound, Optimal Algorithm, and Computationally Efficient Algorithm. Proceedings of the 33rd International Conference on Machine Learning. Ch. 21
- (2022). BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3526113.3545664. Ch. 20
- (2018). Computational Design with Crowds. Computational Interaction. Ch. 16 Ch. 20
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Ch. 20
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Ch. 20
- (2026). Sequential effects in facial attractiveness judgements: No evidence of stable individual differences. Perception. Ch. 16
- (2017). Regret Analysis for Continuous Dueling Bandit. Advances in Neural Information Processing Systems. Ch. 21
- (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. Ch. 17
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 21
- (2025a). Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options. NeurIPS 2025. Ch. 20
- (2022). Gaussian Process Bandit Optimization with Few Batches. International Conference on Artificial Intelligence and Statistics. Ch. 21
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Ch. 16 Ch. 18 Ch. 20
- (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Ch. 20
- (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. Ch. 20
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Ch. 20
- (2016). The appropriacy of averaging in the study of context effects. Psychonomic Bulletin & Review. Ch. 16
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Ch. 19
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Ch. 20
- (2026e). Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium. The Annals of Statistics. doi:10.1214/26-aos2643. Ch. 21
- (1959). Individual Choice Behavior: A Theoretical Analysis. Wiley. Ch. 16 Ch. 20
- (2020). Testing the Random Utility Hypothesis Directly. The Economic Journal. doi:10.1093/ej/uez039. Ch. 16
- (1974). Conditional Logit Analysis of Qualitative Choice Behavior. Frontiers in Econometrics. Ch. 16 Ch. 20
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Ch. 20
- (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. software Ch. 18 Ch. 19
- (2026e). BoTorch CHANGELOG. GitHub. software Ch. 18 Ch. 19
- (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. software Ch. 16 Ch. 18
- (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Ch. 18
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 20
- (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review. Ch. 16
- (2001). Expectation Propagation for Approximate Bayesian Inference. Proceedings of the 17th Conference on Uncertainty in Artificial Intelligence (UAI 2001). Ch. 17
- (2024). Cooperative Multi-Objective Bayesian Design Optimization. ACM Transactions on Interactive Intelligent Systems. doi:10.1145/3657643. Ch. 20
- (2010). Elliptical Slice Sampling. Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS 2010). Ch. 17
- (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Ch. 17 Ch. 20
- (2008). Approximations for Binary Gaussian Process Classification. Journal of Machine Learning Research. Ch. 17
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Ch. 19
- (2017). The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection. Trends in Food Science & Technology. Ch. 16
- (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. software Ch. 18
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Ch. 16 Ch. 19 Ch. 20
- (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. Ch. 20
- (2026). Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs. arXiv. preprint Ch. 20
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Ch. 21
- (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Ch. 20
- (1975). The Analysis of Permutations. Journal of the Royal Statistical Society: Series C (Applied Statistics). Ch. 16 Ch. 20
- (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. Ch. 18
- (2006). Gaussian Processes for Machine Learning. MIT Press. Ch. 17 Ch. 18
- (2021). Optimal Algorithms for Stochastic Contextual Preference Bandits. Advances in Neural Information Processing Systems. Ch. 21
- (2022). Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences. International Conference on Machine Learning. Ch. 21
- (2019b). PAC Battling Bandits in the Plackett-Luce Model. Algorithmic Learning Theory. Ch. 20
- (2021). A Domain-Shrinking based Bayesian Optimization Algorithm with Order-Optimal Regret Performance. Advances in Neural Information Processing Systems. Ch. 21
- (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. Ch. 21
- (2014). When is it Better to Compare than to Score? arXiv. preprint Ch. 16
- (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Ch. 16 Ch. 18
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Ch. 18 Ch. 19
- (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Ch. 17
- (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Ch. 20
- (2020). Scalable Bayesian preference learning for crowds. Machine Learning. Ch. 17 Ch. 20
- (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Ch. 20 Ch. 21
- (2018). When the Good Looks Bad: An Experimental Exploration of the Repulsion Effect. Psychological Science. Ch. 16
- (2021). The elusiveness of context effects in decision making. Trends in Cognitive Sciences. Ch. 16
- (1957). On the Psychophysical Law. Psychological Review. Ch. 16
- (2017a). Correlational Dueling Bandits with Application to Clinical Treatment in Large Decision Spaces. IJCAI 2017. Ch. 21
- (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. Ch. 21
- (2018a). Advancements in Dueling Bandits. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. doi:10.24963/ijcai.2018/776. Ch. 21
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Ch. 17 Ch. 18 Ch. 19
- (1927). A Law of Comparative Judgment. Psychological Review. Ch. 16
- (1986). Accurate Approximations for Posterior Moments and Marginal Densities. Journal of the American Statistical Association. Ch. 17
- (2009). Variational Learning of Inducing Variables in Sparse Gaussian Processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics (AISTATS 2009). Ch. 17
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Ch. 20
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Ch. 20 Ch. 21
- (2013). Generic Exploration and K-armed Voting Bandits. Proceedings of the 30th International Conference on Machine Learning. Ch. 21
- (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 21
- (2021b). Open Problem: Tight Online Confidence Intervals for RKHS Elements. Conference on Learning Theory. Ch. 21
- (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. Ch. 21
- (2019). Decision contamination in the wild: Sequential dependencies in online review ratings. Behavior Research Methods. Ch. 16
- (2023). On the Sublinear Regret of GP-UCB. Advances in Neural Information Processing Systems. Ch. 21
- (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Ch. 19
- (2016). Double Thompson Sampling for Dueling Bandits. Advances in Neural Information Processing Systems. Ch. 21
- (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. preprint Ch. 17 Ch. 20
- (2022). Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation. Quality of Life Research. Ch. 16
- (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. Ch. 21
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 19 Ch. 21
- (1977). The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution. Journal of Mathematical Psychology. Ch. 16
- (2025). Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality. IEEE Transactions on Visualization and Computer Graphics. Ch. 20
- (2009). Interactively optimizing information retrieval systems as a dueling bandits problem. Proceedings of the 26th Annual International Conference on Machine Learning. Ch. 21
- (2012). The K-armed Dueling Bandits Problem. Journal of Computer and System Sciences. Ch. 21
- (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Ch. 20
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Ch. 20
- (2014). Relative Upper Confidence Bound for the K-Armed Dueling Bandit Problem. Proceedings of the 31st International Conference on Machine Learning. Ch. 21
- (2015). Copeland Dueling Bandits. Advances in Neural Information Processing Systems. Ch. 21
- (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Ch. 16