贝叶斯优化
第四部分:从比较中学习
EN

从比较中学习

当目标函数存在于人的头脑中时,最可靠的测量往往是一次比较:选这个,还是选那个。本部分围绕这种测量重建贝叶斯优化。首先说明比较为何有效,并介绍一类已有百年历史的模型,它们把一次选择转化为关于隐藏效用的证据;随后扩展高斯过程,使其能够从比较中学习。此时后验不再是高斯分布,需要借助近似推断。

有了偏好模型,本部分进而构建偏好贝叶斯优化本身:如何选择下一对选项,最优选项期望效用为何是有理论依据的答案,以及人在回路中会带来哪些变化。最后讨论除“两者中哪一个”之外系统还能提出的其他问题,以及支撑这一切的对决赌博机理论。读到本部分中段时,读者本人将成为被优化的对象。

本部分以第二部分与第三部分为前提;第 16 章可以独立阅读。

本部分各章

  1. 16 为什么请人做比较

    测量一个人时,比较为何往往优于评分;哪些模型能把比较转化为关于隐藏效用的证据:心理物理学、Thurstone 的比较判断、Bradley-Terry-Luce 模型、随机效用;一个回答能携带多少信息;以及需要留意的假设。

  2. 17 后验不是高斯分布时

    比较使后验不再是高斯分布。本章以一个效用差为例(其精确后验可以画出),推导并比较 Laplace 近似、期望传播、变分推断与采样;说明精确后验是偏斜正态分布(一般情形下是偏斜高斯过程),偏斜只出现在比较所涉及的方向上;最后报告近似方法的选择有多大影响。

  3. 18 高斯过程偏好学习

    Chu 与 Ghahramani 的模型:效用服从高斯过程,只能通过带噪声的比较来观测,用 Newton 法与 Laplace 近似拟合。本章逐步推导拟合过程,预测新的比较,展示模型在一维、二维及更高维中的表现,说明比较无法识别哪些量、比较图如何影响后验,最后考察 BoTorch 中 PairwiseGP 的实现。

  4. 19 偏好贝叶斯优化

    仅凭对决找到最优选项:对决表述、下一对的选择方法、决策论采集函数 EUBO 及其多选项查询形式 qEUBO、由真人或模拟用户参与的完整循环,以及 2026 年报告的失效模式。

  5. 20 设计提问

    成对比较并非系统唯一可以提出的问题。本章依次讨论多选一与排序、沿直线搜索的滑块、画廊与投影,“差不多”“不确定”“崩溃了”这类回答,多人作答的情形,以及界面为何属于模型。

  6. 21 对决赌博机与比较的理论

    从赌博机的角度讨论如何从对决中学习:偏好不满足传递性时“最优选项”的含义;经典算法及其理论保证;2021 至 2026 年的核化界及其假设与单位;以及至今无人证明的下界。

第四部分参考文献

本部分各章共引用 141 篇文献。

  1. Abeille, M., Faury, L., and Calauzènes, C. (2021). Instance-Wise Minimax-Optimal Algorithms for Logistic Bandits. International Conference on Artificial Intelligence and Statistics. 第 21 章
  2. Alós-Ferrer, C., Fehr, E., and Garagnani, M. (2023). Identifying Nontransitive Preferences. University of Zurich. 工作论文 第 16 章
  3. Apesteguia, J., and Ballester, M. A. (2018). Monotone Stochastic Choice Models: The Case of Risk and Time Preferences. Journal of Political Economy. 第 16 章
  4. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. 第 17 章 第 19 章 第 20 章
  5. Azzalini, A. (1985). A Class of Distributions Which Includes the Normal Ones. Scandinavian Journal of Statistics. 第 17 章
  6. Bagaïni, A., Liu, Y., Kapoor, M., Son, G., Bürkner, P.-C., Tisdall, L., and Mata, R. (2025). A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures. Nature Human Behaviour. doi:10.1038/s41562-024-02085-2. 第 16 章
  7. Balandat, M., Karrer, B., Jiang, D. R., Daulton, S., Letham, B., Wilson, A. G., and Bakshy, E. (2020). BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). 第 19 章
  8. Bavard, S., Lebreton, M., Khamassi, M., Coricelli, G., and Palminteri, S. (2018). Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences. Nature Communications. 第 16 章
  9. Benavoli, A., and Azzimonti, D. (2026a). A tutorial on learning from preferences and choices with Gaussian Processes. Foundations and Trends in Machine Learning 19(1):1-120. 第 20 章
  10. Benavoli, A., Azzimonti, D., and Piga, D. (2021c). Preferential Bayesian optimisation with skew gaussian processes. Proceedings of the Genetic and Evolutionary Computation Conference Companion. 第 17 章 第 20 章
  11. Benavoli, A., Azzimonti, D., and Piga, D. (2023). Learning Choice Functions with Gaussian Processes. Uncertainty in Artificial Intelligence. 第 17 章 第 20 章
  12. Bengs, V., Busa-Fekete, R., El Mesaoudi-Paul, A., and Hüllermeier, E. (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. 第 21 章
  13. Bhatia, S., and Loomes, G. (2017). Noisy preferences in risky choice: A cautionary note. Psychological Review. 第 16 章
  14. Bıyık, E., Palan, M., Landolfi, N. C., Losey, D. P., and Sadigh, D. (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. 第 20 章
  15. Bıyık, E., Huynh, N., Kochenderfer, M. J., and Sadigh, D. (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. 第 17 章 第 18 章
  16. Bradley, R. A., and Terry, M. E. (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika. 第 16 章 第 20 章
  17. Brochu, E., de Freitas, N., and Ghosh, A. (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. 第 18 章 第 19 章 第 20 章
  18. Budish, E., and Kessler, J. B. (2022). Can Market Participants Report Their Preferences Accurately (Enough)? Management Science. 第 16 章
  19. Butler, D. J., and Pogrebna, G. (2018). Predictably intransitive preferences. Judgment and Decision Making. 第 16 章
  20. Chan, L., Liao, Y.-C., Mo, G. B., Dudley, J. J., Cheng, C.-L., Kristensson, P. O., and Oulasvirta, A. (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. 第 19 章
  21. Chang, S., Kim, C.-Y., and Cho, Y. S. (2017). Sequential effects in preference decision: Prior preference assimilates current preference. PLOS ONE. 第 16 章
  22. Chau, S. L., González, J., and Sejdinovic, D. (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. 第 18 章 第 21 章
  23. Chidambaram, K., Seetharaman, K. V., and Syrgkanis, V. (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. 第 20 章
  24. Chong, T. L. H., Shen, I.-C., Sato, I., and Igarashi, T. (2021). Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance. Computer Graphics Forum. doi:10.1111/cgf.14188. 第 20 章
  25. Chowdhury, S. R., and Gopalan, A. (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. 第 21 章
  26. Chu, W., and Ghahramani, Z. (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. 第 16 章 第 17 章 第 18 章 第 19 章 第 21 章
  27. Clark, C. E. (1961). The Greatest of a Finite Set of Random Variables. Operations Research. 第 19 章
  28. Clarke, C. L. A., Vtyurina, A., and Smucker, M. D. (2021). Assessing Top- Preferences. ACM Transactions on Information Systems. 第 16 章
  29. Cover, T. M., and Thomas, J. A. (2006). Elements of Information Theory. Wiley. 第 20 章
  30. Di, Q., He, J., and Gu, Q. (2025). Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback. International Conference on Machine Learning. 第 21 章
  31. Dubey, M., De Peuter, S., Wang, W., and Kaski, S. (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. 预印本 第 20 章
  32. Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M. (2015). Contextual Dueling Bandits. Conference on Learning Theory. 第 21 章
  33. Durante, D. (2019). Conjugate Bayes for probit regression via unified skew-normal distributions. Biometrika. 第 17 章
  34. Enisman, M., Shpitzer, H., and Kleiman, T. (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. 第 16 章
  35. Erarslan, A., Sevilla Salcedo, C., Tanskanen, V., Nisov, A., Päiväkumpu, E., Aisala, H., … Mikkola, P. (2025). Consecutive Preferential Bayesian Optimization. arXiv. 预印本 第 20 章
  36. Faury, L., Abeille, M., Calauzènes, C., and Fercoq, O. (2020). Improved Optimistic Algorithms for Logistic Bandits. International Conference on Machine Learning. 第 21 章
  37. Fauvel, T., and Chalk, M. (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. 预印本 第 19 章
  38. Fechner, G. T. (1860). Elemente der Psychophysik. Breitkopf und Härtel. 第 16 章
  39. Fiedler, M. (1973). Algebraic Connectivity of Graphs. Czechoslovak Mathematical Journal. 第 18 章
  40. Frederick, S., Lee, L., and Baskin, E. (2014). The Limits of Attraction. Journal of Marketing Research. 第 20 章
  41. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. 第 19 章 第 20 章 第 21 章
  42. Heckel, R., Shah, N. B., Ramchandran, K., and Wainwright, M. J. (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. 第 16 章
  43. Hendrickx, J. M., Olshevsky, A., and Saligrama, V. (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. 第 18 章
  44. Hollingworth, H. L. (1910). The Central Tendency of Judgment. The Journal of Philosophy, Psychology and Scientific Methods. 第 16 章
  45. Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M. (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. 预印本 第 16 章 第 18 章
  46. Houlsby, N., Huszár, F., Ghahramani, Z., and Hernández-lobato, J. (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. 第 20 章
  47. Huber, J., Payne, J. W., and Puto, C. (1982). Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis. Journal of Consumer Research. 第 20 章
  48. Jamieson, K., Katariya, S., Deshpande, A., and Nowak, R. (2015). Sparse Dueling Bandits. Proceedings of the 18th International Conference on Artificial Intelligence and Statistics. 第 21 章
  49. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. 第 21 章
  50. Kirschner, J., and Krause, A. (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. 第 21 章
  51. Komiyama, J., Honda, J., Kashima, H., and Nakagawa, H. (2015). Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem. Conference on Learning Theory. 第 21 章
  52. Komiyama, J., Honda, J., and Nakagawa, H. (2016). Copeland Dueling Bandit Problem: Regret Lower Bound, Optimal Algorithm, and Computationally Efficient Algorithm. Proceedings of the 33rd International Conference on Machine Learning. 第 21 章
  53. Koyama, Y., and Goto, M. (2022). BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3526113.3545664. 第 20 章
  54. Koyama, Y., and Igarashi, T. (2018). Computational Design with Crowds. Computational Interaction. 第 16 章 第 20 章
  55. Koyama, Y., Sato, I., Sakamoto, D., and Igarashi, T. (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. 第 20 章
  56. Koyama, Y., Sato, I., and Goto, M. (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). 第 20 章
  57. Kramer, R. S. S., and Cartledge, C. (2026). Sequential effects in facial attractiveness judgements: No evidence of stable individual differences. Perception. 第 16 章
  58. Kumagai, W. (2017). Regret Analysis for Continuous Dueling Bandit. Advances in Neural Information Processing Systems. 第 21 章
  59. Kuss, M., and Rasmussen, C. E. (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. 第 17 章
  60. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. 第 21 章
  61. Lee, J., Yi, S.-w., and Oh, M.-h. (2025a). Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options. NeurIPS 2025. 第 20 章
  62. Li, Z., and Scarlett, J. (2022). Gaussian Process Bandit Optimization with Few Batches. International Conference on Artificial Intelligence and Statistics. 第 21 章
  63. Li, K., Tucker, M., Bıyık, E., Novoseller, E., Burdick, J. W., Sui, Y., … Ames, A. D. (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. 第 16 章 第 18 章 第 20 章
  64. Li, Z., Liao, Y.-C., and Holz, C. (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. 第 20 章
  65. Li, Z., Liao, Y.-C., and Holz, C. (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. 第 20 章
  66. Liao, Y.-C., Belo, J., Moon, H.-S., Steimle, J., and Feit, A. M. (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. 第 20 章
  67. Liew, S. X., Howe, P. D. L., and Little, D. R. (2016). The appropriacy of averaging in the study of context effects. Psychonomic Bulletin & Review. 第 16 章
  68. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. 第 19 章
  69. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. 第 20 章
  70. Liu, K., Long, Q., Shi, Z., Su, W. J., and Xiao, J. (2026e). Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium. The Annals of Statistics. doi:10.1214/26-aos2643. 第 21 章
  71. Luce, R. D. (1959). Individual Choice Behavior: A Theoretical Analysis. Wiley. 第 16 章 第 20 章
  72. McCausland, W. J., Davis-Stober, C., Marley, A., Park, S., and Brown, N. (2020). Testing the Random Utility Hypothesis Directly. The Economic Journal. doi:10.1093/ej/uez039. 第 16 章
  73. McFadden, D. (1974). Conditional Logit Analysis of Qualitative Choice Behavior. Frontiers in Econometrics. 第 16 章 第 20 章
  74. Menn, J., Stenger, D., and Trimpe, S. (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. 第 20 章
  75. Meta Platforms, Inc. (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. 软件 第 18 章 第 19 章
  76. Meta Platforms, Inc. (2026e). BoTorch CHANGELOG. GitHub. 软件 第 18 章 第 19 章
  77. Meta Platforms, Inc. (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. 软件 第 16 章 第 18 章
  78. Meta Platforms, Inc. (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. 软件 第 18 章
  79. Mikkola, P., Todorović, M., Järvi, J., Rinke, P., and Kaski, S. (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. 第 20 章
  80. Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review. 第 16 章
  81. Minka, T. P. (2001). Expectation Propagation for Approximate Bayesian Inference. Proceedings of the 17th Conference on Uncertainty in Artificial Intelligence (UAI 2001). 第 17 章
  82. Mo, G., Dudley, J., Chan, L., Liao, Y.-C., Oulasvirta, A., and Kristensson, P. O. (2024). Cooperative Multi-Objective Bayesian Design Optimization. ACM Transactions on Interactive Intelligent Systems. doi:10.1145/3657643. 第 20 章
  83. Murray, I., Adams, R. P., and MacKay, D. J. C. (2010). Elliptical Slice Sampling. Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS 2010). 第 17 章
  84. Nguyen, Q. P., Tay, S., Low, B. K. H., and Jaillet, P. (2021). Top- Ranking Bayesian Optimization. AAAI 2021. 第 17 章 第 20 章
  85. Nickisch, H., and Rasmussen, C. E. (2008). Approximations for Binary Gaussian Process Classification. Journal of Machine Learning Research. 第 17 章
  86. Niwa, R., Yoshida, S., Koyama, Y., and Ushiku, Y. (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. 第 19 章
  87. O'Mahony, M., and Wichchukit, S. (2017). The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection. Trends in Food Science & Technology. 第 16 章
  88. Optuna developers (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. 软件 第 18 章
  89. Ou, C., Buschek, D., Mayer, S., and Butz, A. (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. 第 16 章 第 19 章 第 20 章
  90. Ou, C., Mayer, S., and Butz, A. (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. 第 20 章
  91. Owaki, T., Koyama, Y., Nakano, T., Yamaguchi, T., Goto, M., and Sakai, H. (2026). Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs. arXiv. 预印本 第 20 章
  92. Pásztor, B., Kassraie, P., and Krause, A. (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. 第 21 章
  93. Peng, Y.-H., Bigham, J. P., and Wu, J. (2026). Efficient Personalization of Generative User Interfaces. arXiv. 预印本 第 20 章
  94. Plackett, R. L. (1975). The Analysis of Permutations. Journal of the Royal Statistical Society: Series C (Applied Statistics). 第 16 章 第 20 章
  95. Pukdee, R., Balcan, M.-F., and Ravikumar, P. (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. 第 18 章
  96. Rasmussen, C. E., and Williams, C. K. I. (2006). Gaussian Processes for Machine Learning. MIT Press. 第 17 章 第 18 章
  97. Saha, A. (2021). Optimal Algorithms for Stochastic Contextual Preference Bandits. Advances in Neural Information Processing Systems. 第 21 章
  98. Saha, A., and Gaillard, P. (2022). Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences. International Conference on Machine Learning. 第 21 章
  99. Saha, A., and Gopalan, A. (2019b). PAC Battling Bandits in the Plackett-Luce Model. Algorithmic Learning Theory. 第 20 章
  100. Salgia, S., Vakili, S., and Zhao, Q. (2021). A Domain-Shrinking based Bayesian Optimization Algorithm with Order-Optimal Regret Performance. Advances in Neural Information Processing Systems. 第 21 章
  101. Scarlett, J., Bogunovic, I., and Cevher, V. (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. 第 21 章
  102. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. (2014). When is it Better to Compare than to Score? arXiv. 预印本 第 16 章
  103. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. J. (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. 第 16 章 第 18 章
  104. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. 预印本 第 18 章 第 19 章
  105. Shvartsman, M., Letham, B., Bakshy, E., and Keeley, S. (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. 第 17 章
  106. Siivola, E., Dhaka, A. K., Andersen, M. R., González, J., García Moreno, P., and Vehtari, A. (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. 第 20 章
  107. Simpson, E., and Gurevych, I. (2020). Scalable Bayesian preference learning for crowds. Machine Learning. 第 17 章 第 20 章
  108. Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. 第 20 章 第 21 章
  109. Spektor, M. S., Kellen, D., and Hotaling, J. M. (2018). When the Good Looks Bad: An Experimental Exploration of the Repulsion Effect. Psychological Science. 第 16 章
  110. Spektor, M. S., Bhatia, S., and Gluth, S. (2021). The elusiveness of context effects in decision making. Trends in Cognitive Sciences. 第 16 章
  111. Stevens, S. S. (1957). On the Psychophysical Law. Psychological Review. 第 16 章
  112. Sui, Y., Yue, Y., and Burdick, J. W. (2017a). Correlational Dueling Bandits with Application to Clinical Treatment in Large Decision Spaces. IJCAI 2017. 第 21 章
  113. Sui, Y., Zhuang, V., Burdick, J. W., and Yue, Y. (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. 第 21 章
  114. Sui, Y., Zoghi, M., Hofmann, K., and Yue, Y. (2018a). Advancements in Dueling Bandits. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. doi:10.24963/ijcai.2018/776. 第 21 章
  115. Takeno, S., Nomura, M., and Karasuyama, M. (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. 第 17 章 第 18 章 第 19 章
  116. Thurstone, L. L. (1927). A Law of Comparative Judgment. Psychological Review. 第 16 章
  117. Tierney, L., and Kadane, J. B. (1986). Accurate Approximations for Posterior Moments and Marginal Densities. Journal of the American Statistical Association. 第 17 章
  118. Titsias, M. (2009). Variational Learning of Inducing Variables in Sparse Gaussian Processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics (AISTATS 2009). 第 17 章
  119. Tucker, M., Cheng, M., Novoseller, E., Cheng, R., Yue, Y., Burdick, J. W., and Ames, A. D. (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. 第 20 章
  120. Tucker, M., Novoseller, E., Kann, C., Sui, Y., Yue, Y., Burdick, J. W., and Ames, A. D. (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). 第 20 章 第 21 章
  121. Urvoy, T., Clerot, F., F\'eraud, R., and Naamane, S. (2013). Generic Exploration and K-armed Voting Bandits. Proceedings of the 30th International Conference on Machine Learning. 第 21 章
  122. Vakili, S., Khezeli, K., and Picheny, V. (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. 第 21 章
  123. Vakili, S., Scarlett, J., and Javidi, T. (2021b). Open Problem: Tight Online Confidence Intervals for RKHS Elements. Conference on Learning Theory. 第 21 章
  124. Verma, A., Dai, Z., Lin, X., Jaillet, P., and Low, B. K. H. (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. 第 21 章
  125. Vinson, D. W., Dale, R., and Jones, M. N. (2019). Decision contamination in the wild: Sequential dependencies in online review ratings. Behavior Research Methods. 第 16 章
  126. Whitehouse, J., Ramdas, A., and Wu, S. (2023). On the Sublinear Regret of GP-UCB. Advances in Neural Information Processing Systems. 第 21 章
  127. Wu, K., and Gardner, J. R. (2026). Knowledge Gradient for Preference Learning. arXiv. 预印本 第 19 章
  128. Wu, H., and Liu, X. (2016). Double Thompson Sampling for Dueling Bandits. Advances in Neural Information Processing Systems. 第 21 章
  129. Wu, K., Sanders, C., Letham, B., and Guan, P. (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. 预印本 第 17 章 第 20 章
  130. Xie, S., Wu, J., and Chen, G. (2022). Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation. Quality of Life Research. 第 16 章
  131. Xu, Y., Joshi, A., Singh, A., and Dubrawski, A. (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. 第 21 章
  132. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. 第 19 章 第 21 章
  133. Yellott, J. J. I. (1977). The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution. Journal of Mathematical Psychology. 第 16 章
  134. Yuan, L.-P., Dudley, J. J., Kristensson, P. O., and Qu, H. (2025). Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality. IEEE Transactions on Visualization and Computer Graphics. 第 20 章
  135. Yue, Y., and Joachims, T. (2009). Interactively optimizing information retrieval systems as a dueling bandits problem. Proceedings of the 26th Annual International Conference on Machine Learning. 第 21 章
  136. Yue, Y., Broder, J., Kleinberg, R., and Joachims, T. (2012). The K-armed Dueling Bandits Problem. Journal of Computer and System Sciences. 第 21 章
  137. Zhang, R., Zhu, X., Pourebadi Khotbehsara, M., Dao, W., Bıyık, E., and Culbertson, H. (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). 第 20 章
  138. Zhu, M., Piga, D., and Bemporad, A. (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. 第 20 章
  139. Zoghi, M., Whiteson, S., Munos, R., and de Rijke, M. (2014). Relative Upper Confidence Bound for the K-Armed Dueling Bandit Problem. Proceedings of the 31st International Conference on Machine Learning. 第 21 章
  140. Zoghi, M., Karnin, Z. S., Whiteson, S., and de Rijke, M. (2015). Copeland Dueling Bandits. Advances in Neural Information Processing Systems. 第 21 章
  141. Zylberberg, A., Bakkour, A., Shohamy, D., and Shadlen, M. N. (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. 第 16 章