Bayesian Optimization
Part IV: Learning from Comparisons
中文

Learning from Comparisons

When the objective lives in a person's head, the most reliable measurement is often a comparison: this one or that one. This part rebuilds Bayesian optimization around that measurement. It starts with why comparisons work and the century-old models that turn a choice into evidence about a hidden utility, then extends the Gaussian process to learn from them, which makes the posterior non-Gaussian and calls for approximate inference.

With a preference model in hand, the part builds preferential Bayesian optimization itself: how to choose the next pair, why the expected utility of the best option is a principled answer, and what a person in the loop changes. It closes with the other questions a system can ask besides "which of two", and with the theory of dueling bandits behind all of it. In the middle of the part, you become the person being optimized.

The part assumes Part II and Part III; Chapter 16 can be read on its own.

Chapters in this part

  1. 16 Why Ask for Comparisons

    Why a comparison is often a better measurement of a person than a rating, and the models that turn one into evidence about a hidden utility: psychophysics, Thurstone's comparative judgment, Bradley-Terry-Luce, random utility, how much one answer can carry, and the assumptions to watch.

  2. 17 When the Posterior Is Not Gaussian

    Comparisons make the posterior non-Gaussian. Using one utility difference whose exact posterior can be drawn, the chapter derives and compares the Laplace approximation, expectation propagation, variational inference, and sampling, shows that the exact answer is a skew-normal (a skew Gaussian process in general) whose skew lives only along compared directions, and reports how much the choice matters.

  3. 18 Gaussian Process Preference Learning

    Chu and Ghahramani's model: a Gaussian process utility observed only through noisy comparisons, fitted by Newton's method with the Laplace approximation. The chapter derives the fit step by step, predicts new comparisons, shows the model in one, two, and more dimensions, explains what comparisons cannot identify and what the comparison graph does to the posterior, and opens BoTorch's PairwiseGP.

  4. 19 Preferential Bayesian Optimization

    Finding the best option from duels alone: the dueling formulation, how to pick the next pair, the decision-theoretic acquisition EUBO and qEUBO, its form for queries of several options, a complete loop with a person or a simulated one, and the failure modes reported in 2026.

  5. 20 Designing the Question

    A pair is not the only question a system can ask. Choices among several and rankings, a slider that searches along a line, galleries and projections, answers that say 'about the same', 'not sure', or 'it crashed', many people at once, and why the interface belongs to the model.

  6. 21 Dueling Bandits and the Theory of Comparisons

    The bandit view of learning from duels: what 'the best option' means when preferences are not transitive, the classic algorithms and their guarantees, the kernelized bounds of 2021 to 2026 with their assumptions and units, and the lower bound nobody has proved.

References for Part IV

141 works cited across this part's chapters.

  1. Abeille, M., Faury, L., and Calauzènes, C. (2021). Instance-Wise Minimax-Optimal Algorithms for Logistic Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 21
  2. Alós-Ferrer, C., Fehr, E., and Garagnani, M. (2023). Identifying Nontransitive Preferences. University of Zurich. working paper Ch. 16
  3. Apesteguia, J., and Ballester, M. A. (2018). Monotone Stochastic Choice Models: The Case of Risk and Time Preferences. Journal of Political Economy. Ch. 16
  4. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Ch. 17 Ch. 19 Ch. 20
  5. Azzalini, A. (1985). A Class of Distributions Which Includes the Normal Ones. Scandinavian Journal of Statistics. Ch. 17
  6. Bagaïni, A., Liu, Y., Kapoor, M., Son, G., Bürkner, P.-C., Tisdall, L., and Mata, R. (2025). A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures. Nature Human Behaviour. doi:10.1038/s41562-024-02085-2. Ch. 16
  7. Balandat, M., Karrer, B., Jiang, D. R., Daulton, S., Letham, B., Wilson, A. G., and Bakshy, E. (2020). BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). Ch. 19
  8. Bavard, S., Lebreton, M., Khamassi, M., Coricelli, G., and Palminteri, S. (2018). Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences. Nature Communications. Ch. 16
  9. Benavoli, A., and Azzimonti, D. (2026a). A tutorial on learning from preferences and choices with Gaussian Processes. Foundations and Trends in Machine Learning 19(1):1-120. Ch. 20
  10. Benavoli, A., Azzimonti, D., and Piga, D. (2021c). Preferential Bayesian optimisation with skew gaussian processes. Proceedings of the Genetic and Evolutionary Computation Conference Companion. Ch. 17 Ch. 20
  11. Benavoli, A., Azzimonti, D., and Piga, D. (2023). Learning Choice Functions with Gaussian Processes. Uncertainty in Artificial Intelligence. Ch. 17 Ch. 20
  12. Bengs, V., Busa-Fekete, R., El Mesaoudi-Paul, A., and Hüllermeier, E. (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. Ch. 21
  13. Bhatia, S., and Loomes, G. (2017). Noisy preferences in risky choice: A cautionary note. Psychological Review. Ch. 16
  14. Bıyık, E., Palan, M., Landolfi, N. C., Losey, D. P., and Sadigh, D. (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Ch. 20
  15. Bıyık, E., Huynh, N., Kochenderfer, M. J., and Sadigh, D. (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. Ch. 17 Ch. 18
  16. Bradley, R. A., and Terry, M. E. (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika. Ch. 16 Ch. 20
  17. Brochu, E., de Freitas, N., and Ghosh, A. (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Ch. 18 Ch. 19 Ch. 20
  18. Budish, E., and Kessler, J. B. (2022). Can Market Participants Report Their Preferences Accurately (Enough)? Management Science. Ch. 16
  19. Butler, D. J., and Pogrebna, G. (2018). Predictably intransitive preferences. Judgment and Decision Making. Ch. 16
  20. Chan, L., Liao, Y.-C., Mo, G. B., Dudley, J. J., Cheng, C.-L., Kristensson, P. O., and Oulasvirta, A. (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. Ch. 19
  21. Chang, S., Kim, C.-Y., and Cho, Y. S. (2017). Sequential effects in preference decision: Prior preference assimilates current preference. PLOS ONE. Ch. 16
  22. Chau, S. L., González, J., and Sejdinovic, D. (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. Ch. 18 Ch. 21
  23. Chidambaram, K., Seetharaman, K. V., and Syrgkanis, V. (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Ch. 20
  24. Chong, T. L. H., Shen, I.-C., Sato, I., and Igarashi, T. (2021). Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance. Computer Graphics Forum. doi:10.1111/cgf.14188. Ch. 20
  25. Chowdhury, S. R., and Gopalan, A. (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. Ch. 21
  26. Chu, W., and Ghahramani, Z. (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Ch. 16 Ch. 17 Ch. 18 Ch. 19 Ch. 21
  27. Clark, C. E. (1961). The Greatest of a Finite Set of Random Variables. Operations Research. Ch. 19
  28. Clarke, C. L. A., Vtyurina, A., and Smucker, M. D. (2021). Assessing Top- Preferences. ACM Transactions on Information Systems. Ch. 16
  29. Cover, T. M., and Thomas, J. A. (2006). Elements of Information Theory. Wiley. Ch. 20
  30. Di, Q., He, J., and Gu, Q. (2025). Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback. International Conference on Machine Learning. Ch. 21
  31. Dubey, M., De Peuter, S., Wang, W., and Kaski, S. (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. preprint Ch. 20
  32. Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M. (2015). Contextual Dueling Bandits. Conference on Learning Theory. Ch. 21
  33. Durante, D. (2019). Conjugate Bayes for probit regression via unified skew-normal distributions. Biometrika. Ch. 17
  34. Enisman, M., Shpitzer, H., and Kleiman, T. (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. Ch. 16
  35. Erarslan, A., Sevilla Salcedo, C., Tanskanen, V., Nisov, A., Päiväkumpu, E., Aisala, H., … Mikkola, P. (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Ch. 20
  36. Faury, L., Abeille, M., Calauzènes, C., and Fercoq, O. (2020). Improved Optimistic Algorithms for Logistic Bandits. International Conference on Machine Learning. Ch. 21
  37. Fauvel, T., and Chalk, M. (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. preprint Ch. 19
  38. Fechner, G. T. (1860). Elemente der Psychophysik. Breitkopf und Härtel. Ch. 16
  39. Fiedler, M. (1973). Algebraic Connectivity of Graphs. Czechoslovak Mathematical Journal. Ch. 18
  40. Frederick, S., Lee, L., and Baskin, E. (2014). The Limits of Attraction. Journal of Marketing Research. Ch. 20
  41. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 19 Ch. 20 Ch. 21
  42. Heckel, R., Shah, N. B., Ramchandran, K., and Wainwright, M. J. (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. Ch. 16
  43. Hendrickx, J. M., Olshevsky, A., and Saligrama, V. (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. Ch. 18
  44. Hollingworth, H. L. (1910). The Central Tendency of Judgment. The Journal of Philosophy, Psychology and Scientific Methods. Ch. 16
  45. Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M. (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Ch. 16 Ch. 18
  46. Houlsby, N., Huszár, F., Ghahramani, Z., and Hernández-lobato, J. (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. Ch. 20
  47. Huber, J., Payne, J. W., and Puto, C. (1982). Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis. Journal of Consumer Research. Ch. 20
  48. Jamieson, K., Katariya, S., Deshpande, A., and Nowak, R. (2015). Sparse Dueling Bandits. Proceedings of the 18th International Conference on Artificial Intelligence and Statistics. Ch. 21
  49. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Ch. 21
  50. Kirschner, J., and Krause, A. (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. Ch. 21
  51. Komiyama, J., Honda, J., Kashima, H., and Nakagawa, H. (2015). Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem. Conference on Learning Theory. Ch. 21
  52. Komiyama, J., Honda, J., and Nakagawa, H. (2016). Copeland Dueling Bandit Problem: Regret Lower Bound, Optimal Algorithm, and Computationally Efficient Algorithm. Proceedings of the 33rd International Conference on Machine Learning. Ch. 21
  53. Koyama, Y., and Goto, M. (2022). BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3526113.3545664. Ch. 20
  54. Koyama, Y., and Igarashi, T. (2018). Computational Design with Crowds. Computational Interaction. Ch. 16 Ch. 20
  55. Koyama, Y., Sato, I., Sakamoto, D., and Igarashi, T. (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Ch. 20
  56. Koyama, Y., Sato, I., and Goto, M. (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Ch. 20
  57. Kramer, R. S. S., and Cartledge, C. (2026). Sequential effects in facial attractiveness judgements: No evidence of stable individual differences. Perception. Ch. 16
  58. Kumagai, W. (2017). Regret Analysis for Continuous Dueling Bandit. Advances in Neural Information Processing Systems. Ch. 21
  59. Kuss, M., and Rasmussen, C. E. (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. Ch. 17
  60. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 21
  61. Lee, J., Yi, S.-w., and Oh, M.-h. (2025a). Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options. NeurIPS 2025. Ch. 20
  62. Li, Z., and Scarlett, J. (2022). Gaussian Process Bandit Optimization with Few Batches. International Conference on Artificial Intelligence and Statistics. Ch. 21
  63. Li, K., Tucker, M., Bıyık, E., Novoseller, E., Burdick, J. W., Sui, Y., … Ames, A. D. (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Ch. 16 Ch. 18 Ch. 20
  64. Li, Z., Liao, Y.-C., and Holz, C. (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Ch. 20
  65. Li, Z., Liao, Y.-C., and Holz, C. (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. Ch. 20
  66. Liao, Y.-C., Belo, J., Moon, H.-S., Steimle, J., and Feit, A. M. (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Ch. 20
  67. Liew, S. X., Howe, P. D. L., and Little, D. R. (2016). The appropriacy of averaging in the study of context effects. Psychonomic Bulletin & Review. Ch. 16
  68. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Ch. 19
  69. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Ch. 20
  70. Liu, K., Long, Q., Shi, Z., Su, W. J., and Xiao, J. (2026e). Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium. The Annals of Statistics. doi:10.1214/26-aos2643. Ch. 21
  71. Luce, R. D. (1959). Individual Choice Behavior: A Theoretical Analysis. Wiley. Ch. 16 Ch. 20
  72. McCausland, W. J., Davis-Stober, C., Marley, A., Park, S., and Brown, N. (2020). Testing the Random Utility Hypothesis Directly. The Economic Journal. doi:10.1093/ej/uez039. Ch. 16
  73. McFadden, D. (1974). Conditional Logit Analysis of Qualitative Choice Behavior. Frontiers in Econometrics. Ch. 16 Ch. 20
  74. Menn, J., Stenger, D., and Trimpe, S. (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Ch. 20
  75. Meta Platforms, Inc. (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. software Ch. 18 Ch. 19
  76. Meta Platforms, Inc. (2026e). BoTorch CHANGELOG. GitHub. software Ch. 18 Ch. 19
  77. Meta Platforms, Inc. (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. software Ch. 16 Ch. 18
  78. Meta Platforms, Inc. (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Ch. 18
  79. Mikkola, P., Todorović, M., Järvi, J., Rinke, P., and Kaski, S. (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 20
  80. Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review. Ch. 16
  81. Minka, T. P. (2001). Expectation Propagation for Approximate Bayesian Inference. Proceedings of the 17th Conference on Uncertainty in Artificial Intelligence (UAI 2001). Ch. 17
  82. Mo, G., Dudley, J., Chan, L., Liao, Y.-C., Oulasvirta, A., and Kristensson, P. O. (2024). Cooperative Multi-Objective Bayesian Design Optimization. ACM Transactions on Interactive Intelligent Systems. doi:10.1145/3657643. Ch. 20
  83. Murray, I., Adams, R. P., and MacKay, D. J. C. (2010). Elliptical Slice Sampling. Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS 2010). Ch. 17
  84. Nguyen, Q. P., Tay, S., Low, B. K. H., and Jaillet, P. (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Ch. 17 Ch. 20
  85. Nickisch, H., and Rasmussen, C. E. (2008). Approximations for Binary Gaussian Process Classification. Journal of Machine Learning Research. Ch. 17
  86. Niwa, R., Yoshida, S., Koyama, Y., and Ushiku, Y. (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Ch. 19
  87. O'Mahony, M., and Wichchukit, S. (2017). The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection. Trends in Food Science & Technology. Ch. 16
  88. Optuna developers (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. software Ch. 18
  89. Ou, C., Buschek, D., Mayer, S., and Butz, A. (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Ch. 16 Ch. 19 Ch. 20
  90. Ou, C., Mayer, S., and Butz, A. (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. Ch. 20
  91. Owaki, T., Koyama, Y., Nakano, T., Yamaguchi, T., Goto, M., and Sakai, H. (2026). Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs. arXiv. preprint Ch. 20
  92. Pásztor, B., Kassraie, P., and Krause, A. (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Ch. 21
  93. Peng, Y.-H., Bigham, J. P., and Wu, J. (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Ch. 20
  94. Plackett, R. L. (1975). The Analysis of Permutations. Journal of the Royal Statistical Society: Series C (Applied Statistics). Ch. 16 Ch. 20
  95. Pukdee, R., Balcan, M.-F., and Ravikumar, P. (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. Ch. 18
  96. Rasmussen, C. E., and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press. Ch. 17 Ch. 18
  97. Saha, A. (2021). Optimal Algorithms for Stochastic Contextual Preference Bandits. Advances in Neural Information Processing Systems. Ch. 21
  98. Saha, A., and Gaillard, P. (2022). Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences. International Conference on Machine Learning. Ch. 21
  99. Saha, A., and Gopalan, A. (2019b). PAC Battling Bandits in the Plackett-Luce Model. Algorithmic Learning Theory. Ch. 20
  100. Salgia, S., Vakili, S., and Zhao, Q. (2021). A Domain-Shrinking based Bayesian Optimization Algorithm with Order-Optimal Regret Performance. Advances in Neural Information Processing Systems. Ch. 21
  101. Scarlett, J., Bogunovic, I., and Cevher, V. (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. Ch. 21
  102. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. (2014). When is it Better to Compare than to Score? arXiv. preprint Ch. 16
  103. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. J. (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Ch. 16 Ch. 18
  104. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Ch. 18 Ch. 19
  105. Shvartsman, M., Letham, B., Bakshy, E., and Keeley, S. (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Ch. 17
  106. Siivola, E., Dhaka, A. K., Andersen, M. R., González, J., García Moreno, P., and Vehtari, A. (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Ch. 20
  107. Simpson, E., and Gurevych, I. (2020). Scalable Bayesian preference learning for crowds. Machine Learning. Ch. 17 Ch. 20
  108. Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Ch. 20 Ch. 21
  109. Spektor, M. S., Kellen, D., and Hotaling, J. M. (2018). When the Good Looks Bad: An Experimental Exploration of the Repulsion Effect. Psychological Science. Ch. 16
  110. Spektor, M. S., Bhatia, S., and Gluth, S. (2021). The elusiveness of context effects in decision making. Trends in Cognitive Sciences. Ch. 16
  111. Stevens, S. S. (1957). On the Psychophysical Law. Psychological Review. Ch. 16
  112. Sui, Y., Yue, Y., and Burdick, J. W. (2017a). Correlational Dueling Bandits with Application to Clinical Treatment in Large Decision Spaces. IJCAI 2017. Ch. 21
  113. Sui, Y., Zhuang, V., Burdick, J. W., and Yue, Y. (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. Ch. 21
  114. Sui, Y., Zoghi, M., Hofmann, K., and Yue, Y. (2018a). Advancements in Dueling Bandits. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. doi:10.24963/ijcai.2018/776. Ch. 21
  115. Takeno, S., Nomura, M., and Karasuyama, M. (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Ch. 17 Ch. 18 Ch. 19
  116. Thurstone, L. L. (1927). A Law of Comparative Judgment. Psychological Review. Ch. 16
  117. Tierney, L., and Kadane, J. B. (1986). Accurate Approximations for Posterior Moments and Marginal Densities. Journal of the American Statistical Association. Ch. 17
  118. Titsias, M. (2009). Variational Learning of Inducing Variables in Sparse Gaussian Processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics (AISTATS 2009). Ch. 17
  119. Tucker, M., Cheng, M., Novoseller, E., Cheng, R., Yue, Y., Burdick, J. W., and Ames, A. D. (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Ch. 20
  120. Tucker, M., Novoseller, E., Kann, C., Sui, Y., Yue, Y., Burdick, J. W., and Ames, A. D. (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Ch. 20 Ch. 21
  121. Urvoy, T., Clerot, F., F\'eraud, R., and Naamane, S. (2013). Generic Exploration and K-armed Voting Bandits. Proceedings of the 30th International Conference on Machine Learning. Ch. 21
  122. Vakili, S., Khezeli, K., and Picheny, V. (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 21
  123. Vakili, S., Scarlett, J., and Javidi, T. (2021b). Open Problem: Tight Online Confidence Intervals for RKHS Elements. Conference on Learning Theory. Ch. 21
  124. Verma, A., Dai, Z., Lin, X., Jaillet, P., and Low, B. K. H. (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. Ch. 21
  125. Vinson, D. W., Dale, R., and Jones, M. N. (2019). Decision contamination in the wild: Sequential dependencies in online review ratings. Behavior Research Methods. Ch. 16
  126. Whitehouse, J., Ramdas, A., and Wu, S. (2023). On the Sublinear Regret of GP-UCB. Advances in Neural Information Processing Systems. Ch. 21
  127. Wu, K., and Gardner, J. R. (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Ch. 19
  128. Wu, H., and Liu, X. (2016). Double Thompson Sampling for Dueling Bandits. Advances in Neural Information Processing Systems. Ch. 21
  129. Wu, K., Sanders, C., Letham, B., and Guan, P. (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. preprint Ch. 17 Ch. 20
  130. Xie, S., Wu, J., and Chen, G. (2022). Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation. Quality of Life Research. Ch. 16
  131. Xu, Y., Joshi, A., Singh, A., and Dubrawski, A. (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. Ch. 21
  132. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 19 Ch. 21
  133. Yellott, J. J. I. (1977). The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution. Journal of Mathematical Psychology. Ch. 16
  134. Yuan, L.-P., Dudley, J. J., Kristensson, P. O., and Qu, H. (2025). Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality. IEEE Transactions on Visualization and Computer Graphics. Ch. 20
  135. Yue, Y., and Joachims, T. (2009). Interactively optimizing information retrieval systems as a dueling bandits problem. Proceedings of the 26th Annual International Conference on Machine Learning. Ch. 21
  136. Yue, Y., Broder, J., Kleinberg, R., and Joachims, T. (2012). The K-armed Dueling Bandits Problem. Journal of Computer and System Sciences. Ch. 21
  137. Zhang, R., Zhu, X., Pourebadi Khotbehsara, M., Dao, W., Bıyık, E., and Culbertson, H. (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Ch. 20
  138. Zhu, M., Piga, D., and Bemporad, A. (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Ch. 20
  139. Zoghi, M., Whiteson, S., Munos, R., and de Rijke, M. (2014). Relative Upper Confidence Bound for the K-Armed Dueling Bandit Problem. Proceedings of the 31st International Conference on Machine Learning. Ch. 21
  140. Zoghi, M., Karnin, Z. S., Whiteson, S., and de Rijke, M. (2015). Copeland Dueling Bandits. Advances in Neural Information Processing Systems. Ch. 21
  141. Zylberberg, A., Bakkour, A., Shohamy, D., and Shadlen, M. N. (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. Ch. 16