研究前沿
前四部分按通常的讲法介绍这些方法。本部分系统梳理截至 2026 年 9 月的文献,报告自 González 等人(2017)以来九年的研究就这些方法确立了哪些结论、哪些仍有争议、哪些尚属缺失。内容涵盖该领域的历史、人的回答背后的模型、选择查询的规则及其有记录的失效、从比较中学习的理论、方法向高维的扩展,以及影响每一项已发表结果的软件与评测实践。
将这几章合起来读,可以看出该领域如今的难点所在。十年间算法日趋成熟:查询的选择有了决策论基础,遗憾界也已追平标量反馈;而描述人如何回答的默认模型,仍停留在 2005 年的形式。瓶颈已经从算法转移到测量:一次比较究竟测量了什么,回答应当如何建模,提问本身又会对回答者产生什么影响(第 45.1 节)。
行文也随之改变。每个论断都附有证据:发表场所、样本量、得出结果的条件,以及是否经过同行评审。本书自身的推断均加以标注。每章最后将结论分为已定、有争议与缺失三类。
本部分以第四部分为前提。
本部分各章
- 26 偏好贝叶斯优化的十年
从 2005 年的基线模型到 2026 年 9 月:该领域如何得名,工具与推断方法如何定型,决策论转向,以及理论迎头赶上、默认流程受到审视的那几年。交互式时间线标出每个里程碑所属的泳道与阶段。
- 27 观测模型、代理模型与推断
似然对人的回答做了哪些假设,哪些代理模型取代了高斯过程及其原因,推断近似的影响有多大,以及默认实现实际做了什么。
- 28 采集函数、查询形式与问题扩展
查询选择规则从启发式到决策理论的演变、多个研究组各自独立发现的失效模式、查询的各种形式与偏好贝叶斯优化的各类问题变体,以及已发表的比较为何不能简单合并:每项比较都在各自的维度与噪声水平下进行。
- 29 理论:从对决赌博机到核化偏好优化
从比较中学习的已证结论:有限臂与线性对决赌博机的结果;2021 至 2026 年的核化遗憾界,及其链接函数、假设与遗憾单位;EUBO 的决策论结果;缺失的下界;观测模型的理论;可识别性;漂移、污染、反应时与停止。
- 30 高维问题与贝叶斯优化格局的变化
贝叶斯优化为何曾有在 10 至 20 维以上失效之说;标量贝叶斯优化在长度尺度先验上得到了什么认识,原因何在;局部偏好方法能扩展到多高的维度,又受哪些混杂因素影响;预训练代理模型、语言模型与成本感知停止用于比较反馈时的现状。
- 31 软件、评测方法与研究社区
偏好贝叶斯优化仍在维护的软件及其默认设置;研究代码为何难以重新运行;方法如何借助模拟用户评测,度量的选择为何决定胜负;换成真人回答时有何变化;以及这一领域由谁研究、分布在哪些学科、研究数量有多少。
第六部分参考文献
本部分各章共引用 271 篇文献。
- (2022). PPBO. GitHub. 软件 第 31 章
- (2019). Multi-objective Bayesian optimisation with preferences over objectives. Advances in Neural Information Processing Systems. 第 28 章
- (2021). Instance-Wise Minimax-Optimal Algorithms for Logistic Bandits. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2025). Bayesian Optimization for Building Social-Influence-Free Consensus. arXiv. 预印本 第 28 章
- (2021). Stochastic Dueling Bandits with Adversarial Corruption. Algorithmic Learning Theory. 第 29 章
- (2022). Batched Dueling Bandits. International Conference on Machine Learning. 第 29 章
- (2026). Best Policy Learning From Trajectory Preference Feedback. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2026). Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences. arXiv. 预印本 第 29 章
- (2026a). Abstract search: preference terms AND "Bayesian optimization". arXiv API. 非同行评审 第 31 章
- (2026b). Abstract search: preferential AND Bayesian AND (optimization OR optimisation). arXiv API. 非同行评审 第 26 章 第 31 章
- (2023a). qEUBO. GitHub. 软件 第 31 章
- (2023b). qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py). GitHub. 软件 第 28 章 第 31 章
- (2020). Multi-attribute Bayesian optimization with interactive preference learning. International Conference on Artificial Intelligence and Statistics. 第 28 章 第 31 章
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. 第 26 章 第 27 章 第 28 章 第 29 章 第 30 章 第 31 章
- (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. 第 27 章 第 28 章 第 31 章
- (2022). Exploiting Composite Functions in Bayesian Optimization. Cornell University. 学位论文 第 31 章
- (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). 第 26 章 第 28 章
- (2026). smac 2.4.1. PyPI. 软件 第 31 章
- (2023). GLIS. GitHub. 软件 第 31 章
- (2021). Global optimization based on active preference learning with radial basis functions. Machine Learning. 第 26 章 第 27 章 第 31 章
- (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. 第 26 章 第 29 章 第 31 章
- (2022). Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity Models. International Conference on Machine Learning. 第 29 章
- (2024). Identifying Copeland Winners in Dueling Bandits with Indifferences. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2026). Time is Knowledge: What Response Times Reveal. working paper (arXiv). 工作论文 第 29 章
- (2026). Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors. arXiv. 预印本 第 30 章
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. 第 26 章 第 27 章 第 28 章 第 30 章
- (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. 第 27 章
- (2024). Dueling Optimization with a Monotone Adversary. International Conference on Algorithmic Learning Theory. 第 29 章
- (2016). Time-Varying Gaussian Process Bandit Optimization. AISTATS 2016. 第 29 章
- (2020). Corruption-Tolerant Gaussian Process Bandit Optimization. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. 第 26 章 第 27 章 第 28 章
- (2010). A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning. arXiv preprint. 预印本 第 31 章
- (2024). Robust Reinforcement Learning from Corrupted Human Feedback. Advances in Neural Information Processing Systems. 第 29 章
- (2021). On Lower Bounds for Standard and Robust Gaussian Process Bandit Optimization. International Conference on Machine Learning. 第 29 章
- (2026). Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries. UAI 2026. 第 29 章
- (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. 第 26 章
- (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. 第 26 章 第 31 章
- (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. 第 27 章 第 29 章
- (2017). Dueling Bandits with Weak Regret. International Conference on Machine Learning. 第 29 章
- (2022). Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation. International Conference on Machine Learning. 第 29 章
- (2026). Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds. Conference on Uncertainty in Artificial Intelligence. 第 28 章
- (2020). Preference-Based Bayesian Optimization in High Dimensions with Human Feedback. SCMLS 2020 Workshop. 研讨会论文 第 28 章
- (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. 第 29 章
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. 第 26 章 第 27 章
- (2020). Human Strategic Steering Improves Performance of Interactive Optimization. UMAP 2020. 第 31 章
- (2022). Safety-Aware Preference-Based Learning for Safety-Critical Control. Learning for Dynamics and Control Conference. 第 28 章
- (2025). Accelerated controller tuning using human feedback and Multi-Task Preferential Bayesian Optimization. 2025 American Control Conference (ACC). 第 28 章
- (2026). Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization. Control Engineering Practice. 第 28 章
- (2023). preferentialBO. GitHub. 软件 第 31 章
- (2025). Experience in Engineering Complex Systems: Active Preference Learning With Multiple Outcomes and Certainty Levels. IEEE Transactions on Human-Machine Systems. 第 27 章
- (2024). Preference Learning of Latent Decision Utilities with a Human-like Model of Preferential Choice. Advances in Neural Information Processing Systems. 第 29 章
- (2024). Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits. International Conference on Learning Representations. 第 29 章
- (2025). Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback. International Conference on Machine Learning. 第 29 章
- (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). 第 26 章 第 30 章
- (2025). Towards Theoretical Understanding of Sequential Decision Making with Preference Feedback. International Conference on Machine Learning. 第 29 章
- (2022). dragonfly-opt 0.1.7. PyPI. 软件 第 31 章
- (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. 预印本 第 27 章
- (2015). Contextual Dueling Bandits. Conference on Learning Theory. 第 29 章
- (2019). Conjugate Bayes for probit regression via unified skew-normal distributions. Biometrika. 第 29 章
- (2024). Efficient Exploration for LLMs. ICML 2024. 第 26 章
- (2026). preferential_batch_bayesian_optimization example. GitHub. 软件 第 31 章
- (2025). Consecutive Preferential Bayesian Optimization. arXiv. 预印本 第 27 章 第 28 章
- (2022). ax-platform 0.2.6. PyPI. 软件 第 26 章 第 31 章
- (2026). Adaptive Candidate Point Thompson Sampling for High-Dimensional Bayesian Optimization. AISTATS 2026. 第 30 章
- (2020). Improved Optimistic Algorithms for Logistic Bandits. International Conference on Machine Learning. 第 29 章
- (2021). Human-in-the-loop optimization of retinal prostheses encoders. Sorbonne Université. 学位论文 第 31 章
- (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. 预印本 第 27 章 第 28 章 第 31 章
- (2017). candy-power-ranking data. GitHub. 非同行评审 第 31 章
- (2018). A Tutorial on Bayesian Optimization. arXiv. 预印本 第 30 章 第 31 章
- (2014). Bayesian Optimization with Inequality Constraints. Proceedings of the 31st International Conference on Machine Learning (ICML 2014). 第 28 章
- (2023). Bayesian Optimization. Cambridge University Press. 第 31 章
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. 第 26 章 第 27 章 第 28 章 第 29 章 第 31 章
- (2026). gpflow 2.11.1. PyPI. 软件 第 31 章
- (2026). gpytorch 1.15.2. PyPI. 软件 第 31 章
- (2023). Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. NeurIPS 2023. 第 27 章 第 28 章 第 30 章 第 31 章
- (2025). LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Findings of the Association for Computational Linguistics: EMNLP 2025. 第 30 章
- (2021a). Identification of the Generalized Condorcet Winner in Multi-dueling Bandits. Advances in Neural Information Processing Systems. 第 29 章
- (2021b). Testification of Condorcet Winners in dueling bandits. Uncertainty in Artificial Intelligence. 第 29 章
- (2026). Elicitation-Augmented Bayesian Optimization. arXiv. 预印本 第 28 章
- (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. 预印本 第 27 章 第 28 章
- (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. 第 27 章
- (2024). HEBO 0.3.6. PyPI. 软件 第 31 章
- (2025). Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization. Journal of Multi-Criteria Decision Analysis. 第 28 章
- (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. 第 26 章 第 27 章 第 30 章
- (2025). Informed Initialization for Bayesian Optimization and Active Learning. NeurIPS 2025. 第 30 章
- (2026). Pitfalls and Remedies for Multi-Task Bayesian Optimization. arXiv. 预印本 第 30 章
- (2023). The Many Facets of Preference-Based Learning. ICML 2023 workshop page. 非同行评审 第 26 章 第 31 章
- (2025). On preference learning based on sequential Bayesian optimization with pairwise comparison. Artificial Intelligence. 第 28 章
- (2026). crashpbo. GitHub. 软件 第 31 章
- (2025). User Preference Meets Pareto-Optimality in Multi-Objective Bayesian Optimization. Proceedings of the AAAI Conference on Artificial Intelligence. 第 28 章
- (2023). A stopping criterion for Bayesian optimization by the gap of expected minimum simple regrets. International Conference on Artificial Intelligence and Statistics. 第 30 章
- (2025). Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. 第 28 章 第 31 章
- (2025). Near-Optimal Algorithm for Non-Stationary Kernelized Bandits. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2026). SUSHI Preference Data Sets. kamishima.net. 非同行评审 第 31 章
- (2025). BOHF_code_submission. GitHub. 软件 第 31 章
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. 第 26 章 第 27 章 第 28 章 第 29 章 第 30 章 第 31 章
- (2025). Efficient Contextual Preferential Bayesian Optimization with Historical Examples. Proceedings of the Genetic and Evolutionary Computation Conference Companion. 第 28 章
- (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. 第 26 章 第 28 章 第 29 章 第 31 章
- (2023). ANACONDA: An Improved Dynamic Regret Algorithm for Adaptive Non-Stationary Dueling Bandits. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. 第 26 章 第 28 章 第 30 章 第 31 章
- (2022). Non-Stationary Dueling Bandits. arXiv. 预印本 第 29 章
- (2015). Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem. Conference on Learning Theory. 第 29 章
- (2017). Computational Design Driven by Visual Aesthetic Preference. The University of Tokyo. doi:10.15083/00076184. 学位论文 第 31 章
- (2025a). preference-regressor.hpp. GitHub. 软件 第 31 章
- (2025b). sequential-line-search. GitHub. 软件 第 31 章
- (2018). Computational Design with Crowds. Computational Interaction. 第 31 章
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. 第 26 章 第 27 章 第 28 章 第 30 章 第 31 章
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). 第 26 章 第 27 章 第 28 章 第 30 章 第 31 章
- (2017). Regret Analysis for Continuous Dueling Bandit. Advances in Neural Information Processing Systems. 第 26 章 第 29 章
- (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. 第 27 章
- (2022). Physically Consistent Preferential Bayesian Optimization for Food Arrangement. IEEE Robotics and Automation Letters. 第 28 章
- (2026). Eliciting Truthful Feedback for Preference-Based Learning via the VCG Mechanism. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2026). Cost-Aware Bayesian Optimization for Prototyping Interactive Devices. CHI 2026. 第 26 章
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. 第 26 章 第 28 章 第 29 章 第 31 章
- (2025a). Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options. NeurIPS 2025. 第 29 章
- (2025). DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization. arXiv. 预印本 第 27 章
- (2022). Gaussian Process Bandit Optimization with Few Batches. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. 第 27 章 第 28 章
- (2022). Detecting Abrupt Changes in Sequential Pairwise Comparison Data. Advances in Neural Information Processing Systems. 第 29 章
- (2024a). Enhancing Preference-based Linear Bandits via Human Response Time. Advances in Neural Information Processing Systems. 第 27 章 第 29 章
- (2024b). Feel-Good Thompson Sampling for Contextual Dueling Bandits. International Conference on Machine Learning. 第 29 章
- (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. 第 26 章 第 27 章 第 28 章
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. 第 26 章
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. 第 26 章 第 28 章 第 31 章
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. 第 27 章 第 28 章 第 30 章
- (2026c). Online Learning and Equilibrium Computation with Ranking Feedback. ICLR 2026. 第 29 章
- (2026e). Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium. The Annals of Statistics. doi:10.1214/26-aos2643. 第 29 章
- (2025). Corruption Robust Offline Reinforcement Learning with Human Feedback. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2024). Bandits with Ranking Feedback. Advances in Neural Information Processing Systems. 第 29 章
- (2019). Sampling Humans for Optimizing Preferences in Coloring Artwork. ICML 2019 Workshop on Human in the Loop Learning. 研讨会论文 第 31 章
- (2025). ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization. arXiv. 预印本 第 30 章
- (2026a). Local Preferential Bayesian Optimization. arXiv. 预印本 第 26 章 第 27 章 第 28 章 第 30 章 第 31 章
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. 第 26 章 第 27 章 第 28 章 第 31 章
- (2026). AEPsych. GitHub. 软件 第 31 章
- (2026a). ax-platform release history. PyPI. 软件 第 31 章
- (2026b). ax/generation_strategy/transition_criterion.py. GitHub. 软件 第 31 章
- (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. 软件 第 28 章 第 31 章
- (2026d). Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1). botorch.org. 软件 第 28 章
- (2026e). BoTorch CHANGELOG. GitHub. 软件 第 26 章 第 27 章 第 28 章 第 30 章 第 31 章
- (2026f). BoTorch LICENSE. GitHub. 软件 第 31 章
- (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. 软件 第 27 章
- (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. 软件 第 27 章 第 30 章 第 31 章
- (2026i). botorch release history. PyPI. 软件 第 31 章
- (2026j). botorch/acquisition/preference.py. GitHub. 软件 第 31 章
- (2026k). botorch/models/utils/gpytorch_modules.py. GitHub. 软件 第 30 章 第 31 章
- (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. 软件 第 26 章 第 31 章
- (2026m). tutorials directory. GitHub. 软件 第 31 章
- (2023). qEUBO. GitHub. 软件 第 31 章
- (2026). lilo. GitHub. 软件 第 31 章
- (2024). Humans as Information Sources in Bayesian Optimization. Aalto University. 学位论文 第 31 章
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. 第 26 章 第 27 章 第 28 章 第 30 章 第 31 章
- (2025). Position: The Future of Bayesian Prediction Is Prior-Fitted. ICML 2025 (position paper). 第 30 章
- (2021). Top- Ranking Bayesian Optimization. AAAI 2021. 第 27 章 第 28 章
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. 第 26 章
- (2021). Online Learning from Human Feedback with Applications to Exoskeleton Gait Optimization. California Institute of Technology. doi:10.7907/gvtx-1586. 学位论文 第 31 章
- (2020). Dueling Posterior Sampling for Preference-Based Reinforcement Learning. Conference on Uncertainty in Artificial Intelligence. 第 29 章
- (2026). Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions. ICML 2026. 第 29 章
- (2026a). Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2025). Ax: A Platform for Adaptive Experimentation. International Conference on Automated Machine Learning. 第 31 章
- (2026a). optuna 5.0.0. PyPI. 软件 第 31 章
- (2026b). optuna-dashboard 0.21.0. PyPI. 软件 第 26 章 第 30 章 第 31 章
- (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. 软件 第 27 章 第 30 章 第 31 章
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. 第 26 章 第 31 章
- (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. 第 26 章 第 30 章 第 31 章
- (2026). PLMBO (Preference Learning Multi-Objective Bayesian Optimization). OptunaHub. 软件 第 31 章
- (2024). Multi-Objective Bayesian Optimization with Active Preference Learning. Proceedings of the AAAI Conference on Artificial Intelligence. 第 28 章 第 31 章
- (2025a). Exploring Exploration in Bayesian Optimization. Conference on Uncertainty in Artificial Intelligence. 第 30 章
- (2025b). Understanding High-Dimensional Bayesian Optimization. ICML 2025, PMLR 267:47902-47923. 第 26 章 第 30 章
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. 第 26 章 第 28 章 第 29 章 第 31 章
- (2025). Towards Uncertainty Unification: A Case Study for Preference Learning. RSS 2025. 第 27 章
- (2026). Machine-generated review of arXiv 2505.23673 (MR-LPF). pith.science. 非同行评审 第 29 章
- (2024). POP-BO. GitHub. 软件 第 31 章
- (2023). GLISp-r: a preference-based optimization algorithm with convergence guarantees. Computational Optimization and Applications. 第 27 章
- (2026). Symposium on Probabilistic Machine Learning website. probml.cc. 非同行评审 第 31 章
- (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. 第 27 章
- (2026). DT-PBO-preprint. GitHub. 软件 第 31 章
- (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. 第 26 章
- (2026). Large language models as uncertainty-calibrated optimizers for experimental discovery. Nature Machine Intelligence. doi:10.1038/s42256-026-01283-z. 第 30 章
- (2026). Zero-shot Bayesian optimization with TabPFN: Competitive with state-of-the-art without per-task training. AutoML Conference 2026 (per Amazon Science page). 第 30 章
- (2024). On Weak Regret Analysis for Dueling Bandits. Advances in Neural Information Processing Systems. 第 29 章
- (2021). Optimal Algorithms for Stochastic Contextual Preference Bandits. Advances in Neural Information Processing Systems. 第 29 章
- (2021). Dueling Bandits with Adversarial Sleeping. Advances in Neural Information Processing Systems. 第 29 章
- (2022). Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences. International Conference on Machine Learning. 第 29 章
- (2019a). Combinatorial Bandits with Relative Feedback. Advances in Neural Information Processing Systems. 第 29 章
- (2019b). PAC Battling Bandits in the Plackett-Luce Model. Algorithmic Learning Theory. 第 29 章
- (2020). From PAC to Instance-Optimal Sample Complexity in the Plackett-Luce Model. International Conference on Machine Learning. 第 29 章
- (2022). Optimal and Efficient Dynamic Regret Algorithms for Non-Stationary Dueling Bandits. International Conference on Machine Learning. 第 29 章
- (2022). Efficient and Optimal Algorithms for Contextual Dueling Bandits under Realizability. International Conference on Algorithmic Learning Theory. 第 29 章
- (2021a). Adversarial Dueling Bandits. International Conference on Machine Learning. 第 29 章
- (2021b). Dueling Convex Optimization. International Conference on Machine Learning. 第 29 章
- (2024). Faster Convergence with MultiWay Preferences. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2025). Dueling Convex Optimization with General Preferences. International Conference on Machine Learning. 第 29 章
- (2021). A Domain-Shrinking based Bayesian Optimization Algorithm with Order-Optimal Regret Performance. Advances in Neural Information Processing Systems. 第 29 章
- (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. 第 29 章
- (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. 第 26 章
- (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. 第 31 章
- (2026). trieste 4.6.0. PyPI. 软件 第 31 章
- (2023). Contextual Bandits and Imitation Learning with Preference-Based Active Queries. Advances in Neural Information Processing Systems. 第 29 章
- (2026). Bulk search: "preferential bayesian optimization". Semantic Scholar API. 非同行评审 第 31 章
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. 预印本 第 26 章 第 27 章 第 28 章 第 31 章
- (2023). GPyOpt (archived). GitHub. 软件 第 31 章
- (2024). Preference-based Pure Exploration. Advances in Neural Information Processing Systems. 第 29 章
- (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. 第 27 章 第 29 章
- (2021). Applications of human feedback in Gaussian processes. Aalto University. 学位论文 第 31 章
- (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. 第 27 章 第 28 章 第 31 章
- (2020). Scalable Bayesian preference learning for crowds. Machine Learning. 第 27 章
- (2026). Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization. Symposium on Probabilistic Machine Learning (ProbML 2026), Proceedings Track. 第 27 章 第 28 章 第 31 章
- (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. 第 29 章
- (2025). Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift. International Conference on Machine Learning. 第 29 章
- (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. 第 26 章 第 28 章 第 29 章
- (2018a). Advancements in Dueling Bandits. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. doi:10.24963/ijcai.2018/776. 第 26 章 第 29 章
- (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. 第 26 章 第 28 章
- (2023). When Can We Track Significant Preference Shifts in Dueling Bandits? Advances in Neural Information Processing Systems. 第 29 章
- (2026). Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis. arXiv. 预印本 第 31 章
- (2022). Preferential Bayesian Optimization with Hallucination Believer. NeurIPS 2022 Workshop on Gaussian Processes, Spatiotemporal Modeling, and Decision-making Systems. 研讨会论文 第 31 章
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. 第 26 章 第 27 章 第 28 章 第 31 章
- (2025). Tackling Biased Evaluators in Dueling Bandits. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-2520. 第 29 章
- (2025). FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization. CHI 2025. 第 27 章
- (2025). Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles. ECC 2025. 第 31 章
- (2026). Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models. European Control Conference. 第 28 章 第 31 章
- (2023). Enabling Robust and User-Customized Bipedal Locomotion on Lower-Body Assistive Devices via Hybrid System Theory and Preference-Based Learning. California Institute of Technology. doi:10.7907/j9hk-xa17. 学位论文 第 31 章
- (2024). POLAR. GitHub. 软件 第 31 章
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. 第 26 章 第 28 章 第 30 章 第 31 章
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). 第 26 章 第 28 章 第 31 章
- (2022). POLAR: Preference Optimization and Learning Algorithms for Robotics. arXiv. 预印本 第 31 章
- (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. 第 29 章
- (2021b). Open Problem: Tight Online Confidence Intervals for RKHS Elements. Conference on Learning Theory. 第 29 章
- (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. 第 27 章 第 29 章
- (2023b). Recent Advances in Bayesian Optimization. ACM Computing Surveys. 第 31 章
- (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. 第 27 章 第 28 章
- (2025b). Fusing Reward and Dueling Feedback in Stochastic Bandits. International Conference on Machine Learning. 第 28 章
- (2025d). Personalized Building Climate Control with Contextual Preferential Bayesian Optimization. arXiv. 预印本 第 28 章
- (2023). On the Sublinear Regret of GP-UCB. Advances in Neural Information Processing Systems. 第 29 章
- (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. 第 26 章
- (2024). Stopping Bayesian Optimization with Probabilistic Regret Bounds. NeurIPS 2024. 第 30 章
- (2026). Knowledge Gradient for Preference Learning. arXiv. 预印本 第 26 章 第 28 章 第 29 章 第 31 章
- (2024). Borda Regret Minimization for Generalized Linear Dueling Bandits. International Conference on Machine Learning. 第 29 章
- (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. 预印本 第 27 章
- (2024). Cost-aware Bayesian Optimization via the Pandora's Box Gittins Index. NeurIPS 2024. 第 30 章
- (2026). Cost-aware Stopping for Bayesian Optimization. International Conference on Machine Learning. 第 30 章
- (2025). Bayesian Optimization with Constraints, Structure and Human Feedback. École Polytechnique Fédérale de Lausanne (EPFL). doi:10.5075/epfl-thesis-11166. 学位论文 第 31 章
- (2020a). Preference-based Reinforcement Learning with Finite-Time Guarantees. Advances in Neural Information Processing Systems. 第 29 章
- (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. 第 28 章 第 29 章
- (2024a). Principled Bayesian Optimisation in Collaboration with Human Experts. NeurIPS 2024. 第 30 章
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. 第 26 章 第 27 章 第 28 章 第 29 章 第 31 章
- (2025b). Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization. ICLR 2025 (oral). 第 26 章 第 30 章
- (2026). GIT-BO: High-Dimensional Bayesian Optimization with Tabular Foundation Models. International Conference on Learning Representations. 第 30 章
- (2026). Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery. ICLR 2026. 第 30 章
- (2009). Interactively optimizing information retrieval systems as a dueling bandits problem. Proceedings of the 26th Annual International Conference on Machine Learning. 第 26 章 第 29 章
- (2012). The K-armed Dueling Bandits Problem. Journal of Computer and System Sciences. 第 29 章
- (2025). PABBO code repository: evaluation config evaluate.yaml. GitHub. 软件 第 27 章 第 30 章 第 31 章
- (2026). PABBO. GitHub. 软件 第 31 章
- (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. 第 26 章 第 27 章 第 28 章 第 30 章 第 31 章
- (2026a). In-Context Multi-Objective Optimization. International Conference on Learning Representations. 第 30 章
- (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). 第 27 章
- (2024). Global and preference-based optimization using surrogate-based methods. IMT School for Advanced Studies Lucca. doi:10.13118/imtlucca/e-theses/415. 学位论文 第 31 章
- (2025). PWAS. GitHub. 软件 第 31 章
- (2025). Global and Preference-Based Optimization with Mixed Variables Using Piecewise Affine Surrogates. Journal of Optimization Theory and Applications. 第 28 章
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. 第 27 章 第 28 章 第 31 章
- (2023). Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. International Conference on Machine Learning. 第 29 章
- (2024). Bayesian Optimisation with Unknown Hyperparameters: Regret Bounds Logarithmically Closer to Optimal. NeurIPS 2024. 第 30 章