从比较中学习
当目标函数存在于人的头脑中时,最可靠的测量往往是一次比较:选这个,还是选那个。本部分围绕这种测量重建贝叶斯优化。首先说明比较为何有效,并介绍一类已有百年历史的模型,它们把一次选择转化为关于隐藏效用的证据;随后扩展高斯过程,使其能够从比较中学习。此时后验不再是高斯分布,需要借助近似推断。
有了偏好模型,本部分进而构建偏好贝叶斯优化本身:如何选择下一对选项,最优选项期望效用为何是有理论依据的答案,以及人在回路中会带来哪些变化。最后讨论除“两者中哪一个”之外系统还能提出的其他问题,以及支撑这一切的对决赌博机理论。读到本部分中段时,读者本人将成为被优化的对象。
本部分以第二部分与第三部分为前提;第 16 章可以独立阅读。
本部分各章
- 16 为什么请人做比较
测量一个人时,比较为何往往优于评分;哪些模型能把比较转化为关于隐藏效用的证据:心理物理学、Thurstone 的比较判断、Bradley-Terry-Luce 模型、随机效用;一个回答能携带多少信息;以及需要留意的假设。
- 17 后验不是高斯分布时
比较使后验不再是高斯分布。本章以一个效用差为例(其精确后验可以画出),推导并比较 Laplace 近似、期望传播、变分推断与采样;说明精确后验是偏斜正态分布(一般情形下是偏斜高斯过程),偏斜只出现在比较所涉及的方向上;最后报告近似方法的选择有多大影响。
- 18 高斯过程偏好学习
Chu 与 Ghahramani 的模型:效用服从高斯过程,只能通过带噪声的比较来观测,用 Newton 法与 Laplace 近似拟合。本章逐步推导拟合过程,预测新的比较,展示模型在一维、二维及更高维中的表现,说明比较无法识别哪些量、比较图如何影响后验,最后考察 BoTorch 中 PairwiseGP 的实现。
- 19 偏好贝叶斯优化
仅凭对决找到最优选项:对决表述、下一对的选择方法、决策论采集函数 EUBO 及其多选项查询形式 qEUBO、由真人或模拟用户参与的完整循环,以及 2026 年报告的失效模式。
- 20 设计提问
成对比较并非系统唯一可以提出的问题。本章依次讨论多选一与排序、沿直线搜索的滑块、画廊与投影,“差不多”“不确定”“崩溃了”这类回答,多人作答的情形,以及界面为何属于模型。
- 21 对决赌博机与比较的理论
从赌博机的角度讨论如何从对决中学习:偏好不满足传递性时“最优选项”的含义;经典算法及其理论保证;2021 至 2026 年的核化界及其假设与单位;以及至今无人证明的下界。
第四部分参考文献
本部分各章共引用 141 篇文献。
- (2021). Instance-Wise Minimax-Optimal Algorithms for Logistic Bandits. International Conference on Artificial Intelligence and Statistics. 第 21 章
- (2023). Identifying Nontransitive Preferences. University of Zurich. 工作论文 第 16 章
- (2018). Monotone Stochastic Choice Models: The Case of Risk and Time Preferences. Journal of Political Economy. 第 16 章
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. 第 17 章 第 19 章 第 20 章
- (1985). A Class of Distributions Which Includes the Normal Ones. Scandinavian Journal of Statistics. 第 17 章
- (2025). A systematic review and meta-analyses of the temporal stability and convergent validity of risk preference measures. Nature Human Behaviour. doi:10.1038/s41562-024-02085-2. 第 16 章
- (2020). BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in Neural Information Processing Systems 33 (NeurIPS 2020). 第 19 章
- (2018). Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences. Nature Communications. 第 16 章
- (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. 第 21 章
- (2017). Noisy preferences in risky choice: A cautionary note. Psychological Review. 第 16 章
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. 第 20 章
- (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. 第 17 章 第 18 章
- (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika. 第 16 章 第 20 章
- (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. 第 18 章 第 19 章 第 20 章
- (2022). Can Market Participants Report Their Preferences Accurately (Enough)? Management Science. 第 16 章
- (2018). Predictably intransitive preferences. Judgment and Decision Making. 第 16 章
- (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. 第 19 章
- (2017). Sequential effects in preference decision: Prior preference assimilates current preference. PLOS ONE. 第 16 章
- (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. 第 18 章 第 21 章
- (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. 第 20 章
- (2021). Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance. Computer Graphics Forum. doi:10.1111/cgf.14188. 第 20 章
- (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. 第 21 章
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. 第 16 章 第 17 章 第 18 章 第 19 章 第 21 章
- (1961). The Greatest of a Finite Set of Random Variables. Operations Research. 第 19 章
- (2021). Assessing Top- Preferences. ACM Transactions on Information Systems. 第 16 章
- (2006). Elements of Information Theory. Wiley. 第 20 章
- (2025). Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback. International Conference on Machine Learning. 第 21 章
- (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. 预印本 第 20 章
- (2015). Contextual Dueling Bandits. Conference on Learning Theory. 第 21 章
- (2019). Conjugate Bayes for probit regression via unified skew-normal distributions. Biometrika. 第 17 章
- (2021). Choice changes preferences, not merely reflects them: A meta-analysis of the artifact-free free-choice paradigm. Journal of Personality and Social Psychology. 第 16 章
- (2025). Consecutive Preferential Bayesian Optimization. arXiv. 预印本 第 20 章
- (2020). Improved Optimistic Algorithms for Logistic Bandits. International Conference on Machine Learning. 第 21 章
- (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. 预印本 第 19 章
- (1860). Elemente der Psychophysik. Breitkopf und Härtel. 第 16 章
- (1973). Algebraic Connectivity of Graphs. Czechoslovak Mathematical Journal. 第 18 章
- (2014). The Limits of Attraction. Journal of Marketing Research. 第 20 章
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. 第 19 章 第 20 章 第 21 章
- (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. 第 16 章
- (2019). Graph Resistance and Learning from Pairwise Comparisons. ICML. 第 18 章
- (1910). The Central Tendency of Judgment. The Journal of Philosophy, Psychology and Scientific Methods. 第 16 章
- (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. 预印本 第 16 章 第 18 章
- (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. 第 20 章
- (1982). Adding Asymmetrically Dominated Alternatives: Violations of Regularity and the Similarity Hypothesis. Journal of Consumer Research. 第 20 章
- (2015). Sparse Dueling Bandits. Proceedings of the 18th International Conference on Artificial Intelligence and Statistics. 第 21 章
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. 第 21 章
- (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. 第 21 章
- (2015). Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem. Conference on Learning Theory. 第 21 章
- (2016). Copeland Dueling Bandit Problem: Regret Lower Bound, Optimal Algorithm, and Computationally Efficient Algorithm. Proceedings of the 33rd International Conference on Machine Learning. 第 21 章
- (2022). BO as Assistant: Using Bayesian Optimization for Asynchronously Generating Design Suggestions. Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3526113.3545664. 第 20 章
- (2018). Computational Design with Crowds. Computational Interaction. 第 16 章 第 20 章
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. 第 20 章
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). 第 20 章
- (2026). Sequential effects in facial attractiveness judgements: No evidence of stable individual differences. Perception. 第 16 章
- (2017). Regret Analysis for Continuous Dueling Bandit. Advances in Neural Information Processing Systems. 第 21 章
- (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. 第 17 章
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. 第 21 章
- (2025a). Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options. NeurIPS 2025. 第 20 章
- (2022). Gaussian Process Bandit Optimization with Few Batches. International Conference on Artificial Intelligence and Statistics. 第 21 章
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. 第 16 章 第 18 章 第 20 章
- (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. 第 20 章
- (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. 第 20 章
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. 第 20 章
- (2016). The appropriacy of averaging in the study of context effects. Psychonomic Bulletin & Review. 第 16 章
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. 第 19 章
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. 第 20 章
- (2026e). Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium. The Annals of Statistics. doi:10.1214/26-aos2643. 第 21 章
- (1959). Individual Choice Behavior: A Theoretical Analysis. Wiley. 第 16 章 第 20 章
- (2020). Testing the Random Utility Hypothesis Directly. The Economic Journal. doi:10.1093/ej/uez039. 第 16 章
- (1974). Conditional Logit Analysis of Qualitative Choice Behavior. Frontiers in Econometrics. 第 16 章 第 20 章
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. 第 20 章
- (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. 软件 第 18 章 第 19 章
- (2026e). BoTorch CHANGELOG. GitHub. 软件 第 18 章 第 19 章
- (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. 软件 第 16 章 第 18 章
- (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. 软件 第 18 章
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. 第 20 章
- (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review. 第 16 章
- (2001). Expectation Propagation for Approximate Bayesian Inference. Proceedings of the 17th Conference on Uncertainty in Artificial Intelligence (UAI 2001). 第 17 章
- (2024). Cooperative Multi-Objective Bayesian Design Optimization. ACM Transactions on Interactive Intelligent Systems. doi:10.1145/3657643. 第 20 章
- (2010). Elliptical Slice Sampling. Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS 2010). 第 17 章
- (2021). Top- Ranking Bayesian Optimization. AAAI 2021. 第 17 章 第 20 章
- (2008). Approximations for Binary Gaussian Process Classification. Journal of Machine Learning Research. 第 17 章
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. 第 19 章
- (2017). The evolution of paired preference tests from forced choice to the use of ‘No Preference’ options, from preference frequencies to d′ values, from placebo pairs to signal detection. Trends in Food Science & Technology. 第 16 章
- (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. 软件 第 18 章
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. 第 16 章 第 19 章 第 20 章
- (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. 第 20 章
- (2026). Learning Feasibility-Aware Latent Spaces for Preference-Based Exploration of Procedural Automotive Wheel Designs. arXiv. 预印本 第 20 章
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. 第 21 章
- (2026). Efficient Personalization of Generative User Interfaces. arXiv. 预印本 第 20 章
- (1975). The Analysis of Permutations. Journal of the Royal Statistical Society: Series C (Applied Statistics). 第 16 章 第 20 章
- (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. 第 18 章
- (2006). Gaussian Processes for Machine Learning. MIT Press. 第 17 章 第 18 章
- (2021). Optimal Algorithms for Stochastic Contextual Preference Bandits. Advances in Neural Information Processing Systems. 第 21 章
- (2022). Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences. International Conference on Machine Learning. 第 21 章
- (2019b). PAC Battling Bandits in the Plackett-Luce Model. Algorithmic Learning Theory. 第 20 章
- (2021). A Domain-Shrinking based Bayesian Optimization Algorithm with Order-Optimal Regret Performance. Advances in Neural Information Processing Systems. 第 21 章
- (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. 第 21 章
- (2014). When is it Better to Compare than to Score? arXiv. 预印本 第 16 章
- (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. 第 16 章 第 18 章
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. 预印本 第 18 章 第 19 章
- (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. 第 17 章
- (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. 第 20 章
- (2020). Scalable Bayesian preference learning for crowds. Machine Learning. 第 17 章 第 20 章
- (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. 第 20 章 第 21 章
- (2018). When the Good Looks Bad: An Experimental Exploration of the Repulsion Effect. Psychological Science. 第 16 章
- (2021). The elusiveness of context effects in decision making. Trends in Cognitive Sciences. 第 16 章
- (1957). On the Psychophysical Law. Psychological Review. 第 16 章
- (2017a). Correlational Dueling Bandits with Application to Clinical Treatment in Large Decision Spaces. IJCAI 2017. 第 21 章
- (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. 第 21 章
- (2018a). Advancements in Dueling Bandits. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. doi:10.24963/ijcai.2018/776. 第 21 章
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. 第 17 章 第 18 章 第 19 章
- (1927). A Law of Comparative Judgment. Psychological Review. 第 16 章
- (1986). Accurate Approximations for Posterior Moments and Marginal Densities. Journal of the American Statistical Association. 第 17 章
- (2009). Variational Learning of Inducing Variables in Sparse Gaussian Processes. Proceedings of the 12th International Conference on Artificial Intelligence and Statistics (AISTATS 2009). 第 17 章
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. 第 20 章
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). 第 20 章 第 21 章
- (2013). Generic Exploration and K-armed Voting Bandits. Proceedings of the 30th International Conference on Machine Learning. 第 21 章
- (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. 第 21 章
- (2021b). Open Problem: Tight Online Confidence Intervals for RKHS Elements. Conference on Learning Theory. 第 21 章
- (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. 第 21 章
- (2019). Decision contamination in the wild: Sequential dependencies in online review ratings. Behavior Research Methods. 第 16 章
- (2023). On the Sublinear Regret of GP-UCB. Advances in Neural Information Processing Systems. 第 21 章
- (2026). Knowledge Gradient for Preference Learning. arXiv. 预印本 第 19 章
- (2016). Double Thompson Sampling for Dueling Bandits. Advances in Neural Information Processing Systems. 第 21 章
- (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. 预印本 第 17 章 第 20 章
- (2022). Discrete choice experiment with duration versus time trade-off: a comparison of test–retest reliability of health utility elicitation approaches in SF-6Dv2 valuation. Quality of Life Research. 第 16 章
- (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. 第 21 章
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. 第 19 章 第 21 章
- (1977). The relationship between Luce's Choice Axiom, Thurstone's Theory of Comparative Judgment, and the double exponential distribution. Journal of Mathematical Psychology. 第 16 章
- (2025). Personalized Dual-Level Color Grading for 360-degree Images in Virtual Reality. IEEE Transactions on Visualization and Computer Graphics. 第 20 章
- (2009). Interactively optimizing information retrieval systems as a dueling bandits problem. Proceedings of the 26th Annual International Conference on Machine Learning. 第 21 章
- (2012). The K-armed Dueling Bandits Problem. Journal of Computer and System Sciences. 第 21 章
- (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). 第 20 章
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. 第 20 章
- (2014). Relative Upper Confidence Bound for the K-Armed Dueling Bandit Problem. Proceedings of the 31st International Conference on Machine Learning. 第 21 章
- (2015). Copeland Dueling Bandits. Advances in Neural Information Processing Systems. 第 21 章
- (2024). Value construction through sequential sampling explains serial dependencies in decision making. eLife. doi:10.7554/eLife.96997. 第 16 章