The Research Frontier
The first four parts present the methods as they are usually taught. This part reports what nine years of research since González et al. (2017) have established about them, what remains contested, and what is missing, from a systematic reading of the literature through September 2026. It covers the history of the field, the models behind a human answer, the rules for choosing queries and their documented failures, the theory of learning from comparisons, the scaling of the methods to many dimensions, and the software and evaluation practices that shape every published result.
Read together, the chapters show where the field's difficulty now lies. The algorithms matured over the decade, with a decision-theoretic foundation for choosing queries and regret bounds that caught up with scalar feedback, while the default model of a human answer stayed as it was in 2005. The bottleneck has moved from algorithms to measurement: what a single comparison measures, how answers should be modeled, and what asking does to the person who answers (Section 45.1).
The tone changes accordingly. Claims carry their evidence: venues, sample sizes, the conditions under which a result was observed, and whether it has been peer reviewed. The book's own inferences are marked. Each chapter ends by sorting its conclusions into what is settled, what is contested, and what is missing.
The part assumes Part IV.
Chapters in this part
- 26 A Decade of Preferential Bayesian Optimization
From the 2005 baselines to September 2026: how the field got its name, how its tools and inference settled, the decision-theoretic turn, and the years in which theory caught up and the default pipeline came under scrutiny. An interactive timeline places every milestone in its lane and phase.
- 27 Observation Models, Surrogates, and Inference
What the likelihood assumes about a human answer, which surrogates replace the Gaussian process and why, how much the inference approximation matters, and what the default implementation actually does.
- 28 Acquisition, Query Forms, and Problem Extensions
How the rules for choosing queries evolved from heuristics to decision theory, the failure modes several groups found independently, the forms a query can take and the problem variants built on preferential Bayesian optimization, and why the published comparisons, each run at its own dimension and noise level, cannot simply be pooled.
- 29 Theory: From Dueling Bandits to Kernelized Preference Optimization
What is proved about learning from comparisons: the finite-arm and linear dueling-bandit results, the kernelized regret bounds of 2021 to 2026 with their links, assumptions, and regret units, the decision-theoretic results for EUBO, the missing lower bounds, the theory of the observation model, identifiability, and drift, contamination, response times, and stopping.
- 30 High Dimensions and the Changing Landscape of Bayesian Optimization
Why Bayesian optimization was said to fail beyond 10 to 20 dimensions, what scalar BO learned about lengthscale priors and why, how far local preferential methods reach and what confounds them, and where pretrained surrogates, language models, and cost-aware stopping stand for comparisons.
- 31 Software, Evaluation, and the Research Community
The maintained software for PBO and the defaults it ships, why research code is hard to rerun, how methods are evaluated with simulated users and why the choice of metric decides the winner, what changes when real people answer, and who does this research, in which disciplines, and how much of it there is.
References for Part VI
271 works cited across this part's chapters.
- (2022). PPBO. GitHub. software Ch. 31
- (2019). Multi-objective Bayesian optimisation with preferences over objectives. Advances in Neural Information Processing Systems. Ch. 28
- (2021). Instance-Wise Minimax-Optimal Algorithms for Logistic Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2025). Bayesian Optimization for Building Social-Influence-Free Consensus. arXiv. preprint Ch. 28
- (2021). Stochastic Dueling Bandits with Adversarial Corruption. Algorithmic Learning Theory. Ch. 29
- (2022). Batched Dueling Bandits. International Conference on Machine Learning. Ch. 29
- (2026). Best Policy Learning From Trajectory Preference Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2026). Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences. arXiv. preprint Ch. 29
- (2026a). Abstract search: preference terms AND "Bayesian optimization". arXiv API. non-peer-reviewed Ch. 31
- (2026b). Abstract search: preferential AND Bayesian AND (optimization OR optimisation). arXiv API. non-peer-reviewed Ch. 26 Ch. 31
- (2023a). qEUBO. GitHub. software Ch. 31
- (2023b). qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py). GitHub. software Ch. 28 Ch. 31
- (2020). Multi-attribute Bayesian optimization with interactive preference learning. International Conference on Artificial Intelligence and Statistics. Ch. 28 Ch. 31
- (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 30 Ch. 31
- (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Ch. 27 Ch. 28 Ch. 31
- (2022). Exploiting Composite Functions in Bayesian Optimization. Cornell University. thesis Ch. 31
- (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Ch. 26 Ch. 28
- (2026). smac 2.4.1. PyPI. software Ch. 31
- (2023). GLIS. GitHub. software Ch. 31
- (2021). Global optimization based on active preference learning with radial basis functions. Machine Learning. Ch. 26 Ch. 27 Ch. 31
- (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. Ch. 26 Ch. 29 Ch. 31
- (2022). Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity Models. International Conference on Machine Learning. Ch. 29
- (2024). Identifying Copeland Winners in Dueling Bandits with Indifferences. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2026). Time is Knowledge: What Response Times Reveal. working paper (arXiv). working paper Ch. 29
- (2026). Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors. arXiv. preprint Ch. 30
- (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Ch. 26 Ch. 27 Ch. 28 Ch. 30
- (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. Ch. 27
- (2024). Dueling Optimization with a Monotone Adversary. International Conference on Algorithmic Learning Theory. Ch. 29
- (2016). Time-Varying Gaussian Process Bandit Optimization. AISTATS 2016. Ch. 29
- (2020). Corruption-Tolerant Gaussian Process Bandit Optimization. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Ch. 26 Ch. 27 Ch. 28
- (2010). A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning. arXiv preprint. preprint Ch. 31
- (2024). Robust Reinforcement Learning from Corrupted Human Feedback. Advances in Neural Information Processing Systems. Ch. 29
- (2021). On Lower Bounds for Standard and Robust Gaussian Process Bandit Optimization. International Conference on Machine Learning. Ch. 29
- (2026). Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries. UAI 2026. Ch. 29
- (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. Ch. 26
- (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. Ch. 26 Ch. 31
- (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. Ch. 27 Ch. 29
- (2017). Dueling Bandits with Weak Regret. International Conference on Machine Learning. Ch. 29
- (2022). Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation. International Conference on Machine Learning. Ch. 29
- (2026). Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds. Conference on Uncertainty in Artificial Intelligence. Ch. 28
- (2020). Preference-Based Bayesian Optimization in High Dimensions with Human Feedback. SCMLS 2020 Workshop. workshop paper Ch. 28
- (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. Ch. 29
- (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Ch. 26 Ch. 27
- (2020). Human Strategic Steering Improves Performance of Interactive Optimization. UMAP 2020. Ch. 31
- (2022). Safety-Aware Preference-Based Learning for Safety-Critical Control. Learning for Dynamics and Control Conference. Ch. 28
- (2025). Accelerated controller tuning using human feedback and Multi-Task Preferential Bayesian Optimization. 2025 American Control Conference (ACC). Ch. 28
- (2026). Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization. Control Engineering Practice. Ch. 28
- (2023). preferentialBO. GitHub. software Ch. 31
- (2025). Experience in Engineering Complex Systems: Active Preference Learning With Multiple Outcomes and Certainty Levels. IEEE Transactions on Human-Machine Systems. Ch. 27
- (2024). Preference Learning of Latent Decision Utilities with a Human-like Model of Preferential Choice. Advances in Neural Information Processing Systems. Ch. 29
- (2024). Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits. International Conference on Learning Representations. Ch. 29
- (2025). Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback. International Conference on Machine Learning. Ch. 29
- (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). Ch. 26 Ch. 30
- (2025). Towards Theoretical Understanding of Sequential Decision Making with Preference Feedback. International Conference on Machine Learning. Ch. 29
- (2022). dragonfly-opt 0.1.7. PyPI. software Ch. 31
- (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. preprint Ch. 27
- (2015). Contextual Dueling Bandits. Conference on Learning Theory. Ch. 29
- (2019). Conjugate Bayes for probit regression via unified skew-normal distributions. Biometrika. Ch. 29
- (2024). Efficient Exploration for LLMs. ICML 2024. Ch. 26
- (2026). preferential_batch_bayesian_optimization example. GitHub. software Ch. 31
- (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Ch. 27 Ch. 28
- (2022). ax-platform 0.2.6. PyPI. software Ch. 26 Ch. 31
- (2026). Adaptive Candidate Point Thompson Sampling for High-Dimensional Bayesian Optimization. AISTATS 2026. Ch. 30
- (2020). Improved Optimistic Algorithms for Logistic Bandits. International Conference on Machine Learning. Ch. 29
- (2021). Human-in-the-loop optimization of retinal prostheses encoders. Sorbonne Université. thesis Ch. 31
- (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. preprint Ch. 27 Ch. 28 Ch. 31
- (2017). candy-power-ranking data. GitHub. non-peer-reviewed Ch. 31
- (2018). A Tutorial on Bayesian Optimization. arXiv. preprint Ch. 30 Ch. 31
- (2014). Bayesian Optimization with Inequality Constraints. Proceedings of the 31st International Conference on Machine Learning (ICML 2014). Ch. 28
- (2023). Bayesian Optimization. Cambridge University Press. Ch. 31
- (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 31
- (2026). gpflow 2.11.1. PyPI. software Ch. 31
- (2026). gpytorch 1.15.2. PyPI. software Ch. 31
- (2023). Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. NeurIPS 2023. Ch. 27 Ch. 28 Ch. 30 Ch. 31
- (2025). LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Findings of the Association for Computational Linguistics: EMNLP 2025. Ch. 30
- (2021a). Identification of the Generalized Condorcet Winner in Multi-dueling Bandits. Advances in Neural Information Processing Systems. Ch. 29
- (2021b). Testification of Condorcet Winners in dueling bandits. Uncertainty in Artificial Intelligence. Ch. 29
- (2026). Elicitation-Augmented Bayesian Optimization. arXiv. preprint Ch. 28
- (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Ch. 27 Ch. 28
- (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. Ch. 27
- (2024). HEBO 0.3.6. PyPI. software Ch. 31
- (2025). Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization. Journal of Multi-Criteria Decision Analysis. Ch. 28
- (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 30
- (2025). Informed Initialization for Bayesian Optimization and Active Learning. NeurIPS 2025. Ch. 30
- (2026). Pitfalls and Remedies for Multi-Task Bayesian Optimization. arXiv. preprint Ch. 30
- (2023). The Many Facets of Preference-Based Learning. ICML 2023 workshop page. non-peer-reviewed Ch. 26 Ch. 31
- (2025). On preference learning based on sequential Bayesian optimization with pairwise comparison. Artificial Intelligence. Ch. 28
- (2026). crashpbo. GitHub. software Ch. 31
- (2025). User Preference Meets Pareto-Optimality in Multi-Objective Bayesian Optimization. Proceedings of the AAAI Conference on Artificial Intelligence. Ch. 28
- (2023). A stopping criterion for Bayesian optimization by the gap of expected minimum simple regrets. International Conference on Artificial Intelligence and Statistics. Ch. 30
- (2025). Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. Ch. 28 Ch. 31
- (2025). Near-Optimal Algorithm for Non-Stationary Kernelized Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2026). SUSHI Preference Data Sets. kamishima.net. non-peer-reviewed Ch. 31
- (2025). BOHF_code_submission. GitHub. software Ch. 31
- (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 30 Ch. 31
- (2025). Efficient Contextual Preferential Bayesian Optimization with Historical Examples. Proceedings of the Genetic and Evolutionary Computation Conference Companion. Ch. 28
- (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. Ch. 26 Ch. 28 Ch. 29 Ch. 31
- (2023). ANACONDA: An Improved Dynamic Regret Algorithm for Adaptive Non-Stationary Dueling Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Ch. 26 Ch. 28 Ch. 30 Ch. 31
- (2022). Non-Stationary Dueling Bandits. arXiv. preprint Ch. 29
- (2015). Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem. Conference on Learning Theory. Ch. 29
- (2017). Computational Design Driven by Visual Aesthetic Preference. The University of Tokyo. doi:10.15083/00076184. thesis Ch. 31
- (2025a). preference-regressor.hpp. GitHub. software Ch. 31
- (2025b). sequential-line-search. GitHub. software Ch. 31
- (2018). Computational Design with Crowds. Computational Interaction. Ch. 31
- (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
- (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
- (2017). Regret Analysis for Continuous Dueling Bandit. Advances in Neural Information Processing Systems. Ch. 26 Ch. 29
- (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. Ch. 27
- (2022). Physically Consistent Preferential Bayesian Optimization for Food Arrangement. IEEE Robotics and Automation Letters. Ch. 28
- (2026). Eliciting Truthful Feedback for Preference-Based Learning via the VCG Mechanism. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2026). Cost-Aware Bayesian Optimization for Prototyping Interactive Devices. CHI 2026. Ch. 26
- (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 26 Ch. 28 Ch. 29 Ch. 31
- (2025a). Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options. NeurIPS 2025. Ch. 29
- (2025). DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization. arXiv. preprint Ch. 27
- (2022). Gaussian Process Bandit Optimization with Few Batches. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Ch. 27 Ch. 28
- (2022). Detecting Abrupt Changes in Sequential Pairwise Comparison Data. Advances in Neural Information Processing Systems. Ch. 29
- (2024a). Enhancing Preference-based Linear Bandits via Human Response Time. Advances in Neural Information Processing Systems. Ch. 27 Ch. 29
- (2024b). Feel-Good Thompson Sampling for Contextual Dueling Bandits. International Conference on Machine Learning. Ch. 29
- (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Ch. 26 Ch. 27 Ch. 28
- (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Ch. 26
- (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Ch. 26 Ch. 28 Ch. 31
- (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Ch. 27 Ch. 28 Ch. 30
- (2026c). Online Learning and Equilibrium Computation with Ranking Feedback. ICLR 2026. Ch. 29
- (2026e). Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium. The Annals of Statistics. doi:10.1214/26-aos2643. Ch. 29
- (2025). Corruption Robust Offline Reinforcement Learning with Human Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2024). Bandits with Ranking Feedback. Advances in Neural Information Processing Systems. Ch. 29
- (2019). Sampling Humans for Optimizing Preferences in Coloring Artwork. ICML 2019 Workshop on Human in the Loop Learning. workshop paper Ch. 31
- (2025). ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization. arXiv. preprint Ch. 30
- (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
- (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Ch. 26 Ch. 27 Ch. 28 Ch. 31
- (2026). AEPsych. GitHub. software Ch. 31
- (2026a). ax-platform release history. PyPI. software Ch. 31
- (2026b). ax/generation_strategy/transition_criterion.py. GitHub. software Ch. 31
- (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. software Ch. 28 Ch. 31
- (2026d). Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1). botorch.org. software Ch. 28
- (2026e). BoTorch CHANGELOG. GitHub. software Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
- (2026f). BoTorch LICENSE. GitHub. software Ch. 31
- (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. software Ch. 27
- (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Ch. 27 Ch. 30 Ch. 31
- (2026i). botorch release history. PyPI. software Ch. 31
- (2026j). botorch/acquisition/preference.py. GitHub. software Ch. 31
- (2026k). botorch/models/utils/gpytorch_modules.py. GitHub. software Ch. 30 Ch. 31
- (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Ch. 26 Ch. 31
- (2026m). tutorials directory. GitHub. software Ch. 31
- (2023). qEUBO. GitHub. software Ch. 31
- (2026). lilo. GitHub. software Ch. 31
- (2024). Humans as Information Sources in Bayesian Optimization. Aalto University. thesis Ch. 31
- (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
- (2025). Position: The Future of Bayesian Prediction Is Prior-Fitted. ICML 2025 (position paper). Ch. 30
- (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Ch. 27 Ch. 28
- (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Ch. 26
- (2021). Online Learning from Human Feedback with Applications to Exoskeleton Gait Optimization. California Institute of Technology. doi:10.7907/gvtx-1586. thesis Ch. 31
- (2020). Dueling Posterior Sampling for Preference-Based Reinforcement Learning. Conference on Uncertainty in Artificial Intelligence. Ch. 29
- (2026). Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions. ICML 2026. Ch. 29
- (2026a). Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2025). Ax: A Platform for Adaptive Experimentation. International Conference on Automated Machine Learning. Ch. 31
- (2026a). optuna 5.0.0. PyPI. software Ch. 31
- (2026b). optuna-dashboard 0.21.0. PyPI. software Ch. 26 Ch. 30 Ch. 31
- (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. software Ch. 27 Ch. 30 Ch. 31
- (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Ch. 26 Ch. 31
- (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. Ch. 26 Ch. 30 Ch. 31
- (2026). PLMBO (Preference Learning Multi-Objective Bayesian Optimization). OptunaHub. software Ch. 31
- (2024). Multi-Objective Bayesian Optimization with Active Preference Learning. Proceedings of the AAAI Conference on Artificial Intelligence. Ch. 28 Ch. 31
- (2025a). Exploring Exploration in Bayesian Optimization. Conference on Uncertainty in Artificial Intelligence. Ch. 30
- (2025b). Understanding High-Dimensional Bayesian Optimization. ICML 2025, PMLR 267:47902-47923. Ch. 26 Ch. 30
- (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Ch. 26 Ch. 28 Ch. 29 Ch. 31
- (2025). Towards Uncertainty Unification: A Case Study for Preference Learning. RSS 2025. Ch. 27
- (2026). Machine-generated review of arXiv 2505.23673 (MR-LPF). pith.science. non-peer-reviewed Ch. 29
- (2024). POP-BO. GitHub. software Ch. 31
- (2023). GLISp-r: a preference-based optimization algorithm with convergence guarantees. Computational Optimization and Applications. Ch. 27
- (2026). Symposium on Probabilistic Machine Learning website. probml.cc. non-peer-reviewed Ch. 31
- (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. Ch. 27
- (2026). DT-PBO-preprint. GitHub. software Ch. 31
- (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. Ch. 26
- (2026). Large language models as uncertainty-calibrated optimizers for experimental discovery. Nature Machine Intelligence. doi:10.1038/s42256-026-01283-z. Ch. 30
- (2026). Zero-shot Bayesian optimization with TabPFN: Competitive with state-of-the-art without per-task training. AutoML Conference 2026 (per Amazon Science page). Ch. 30
- (2024). On Weak Regret Analysis for Dueling Bandits. Advances in Neural Information Processing Systems. Ch. 29
- (2021). Optimal Algorithms for Stochastic Contextual Preference Bandits. Advances in Neural Information Processing Systems. Ch. 29
- (2021). Dueling Bandits with Adversarial Sleeping. Advances in Neural Information Processing Systems. Ch. 29
- (2022). Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences. International Conference on Machine Learning. Ch. 29
- (2019a). Combinatorial Bandits with Relative Feedback. Advances in Neural Information Processing Systems. Ch. 29
- (2019b). PAC Battling Bandits in the Plackett-Luce Model. Algorithmic Learning Theory. Ch. 29
- (2020). From PAC to Instance-Optimal Sample Complexity in the Plackett-Luce Model. International Conference on Machine Learning. Ch. 29
- (2022). Optimal and Efficient Dynamic Regret Algorithms for Non-Stationary Dueling Bandits. International Conference on Machine Learning. Ch. 29
- (2022). Efficient and Optimal Algorithms for Contextual Dueling Bandits under Realizability. International Conference on Algorithmic Learning Theory. Ch. 29
- (2021a). Adversarial Dueling Bandits. International Conference on Machine Learning. Ch. 29
- (2021b). Dueling Convex Optimization. International Conference on Machine Learning. Ch. 29
- (2024). Faster Convergence with MultiWay Preferences. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2025). Dueling Convex Optimization with General Preferences. International Conference on Machine Learning. Ch. 29
- (2021). A Domain-Shrinking based Bayesian Optimization Algorithm with Order-Optimal Regret Performance. Advances in Neural Information Processing Systems. Ch. 29
- (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. Ch. 29
- (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Ch. 26
- (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Ch. 31
- (2026). trieste 4.6.0. PyPI. software Ch. 31
- (2023). Contextual Bandits and Imitation Learning with Preference-Based Active Queries. Advances in Neural Information Processing Systems. Ch. 29
- (2026). Bulk search: "preferential bayesian optimization". Semantic Scholar API. non-peer-reviewed Ch. 31
- (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Ch. 26 Ch. 27 Ch. 28 Ch. 31
- (2023). GPyOpt (archived). GitHub. software Ch. 31
- (2024). Preference-based Pure Exploration. Advances in Neural Information Processing Systems. Ch. 29
- (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Ch. 27 Ch. 29
- (2021). Applications of human feedback in Gaussian processes. Aalto University. thesis Ch. 31
- (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Ch. 27 Ch. 28 Ch. 31
- (2020). Scalable Bayesian preference learning for crowds. Machine Learning. Ch. 27
- (2026). Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization. Symposium on Probabilistic Machine Learning (ProbML 2026), Proceedings Track. Ch. 27 Ch. 28 Ch. 31
- (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Ch. 29
- (2025). Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift. International Conference on Machine Learning. Ch. 29
- (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. Ch. 26 Ch. 28 Ch. 29
- (2018a). Advancements in Dueling Bandits. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. doi:10.24963/ijcai.2018/776. Ch. 26 Ch. 29
- (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. Ch. 26 Ch. 28
- (2023). When Can We Track Significant Preference Shifts in Dueling Bandits? Advances in Neural Information Processing Systems. Ch. 29
- (2026). Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis. arXiv. preprint Ch. 31
- (2022). Preferential Bayesian Optimization with Hallucination Believer. NeurIPS 2022 Workshop on Gaussian Processes, Spatiotemporal Modeling, and Decision-making Systems. workshop paper Ch. 31
- (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 31
- (2025). Tackling Biased Evaluators in Dueling Bandits. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-2520. Ch. 29
- (2025). FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization. CHI 2025. Ch. 27
- (2025). Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles. ECC 2025. Ch. 31
- (2026). Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models. European Control Conference. Ch. 28 Ch. 31
- (2023). Enabling Robust and User-Customized Bipedal Locomotion on Lower-Body Assistive Devices via Hybrid System Theory and Preference-Based Learning. California Institute of Technology. doi:10.7907/j9hk-xa17. thesis Ch. 31
- (2024). POLAR. GitHub. software Ch. 31
- (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Ch. 26 Ch. 28 Ch. 30 Ch. 31
- (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Ch. 26 Ch. 28 Ch. 31
- (2022). POLAR: Preference Optimization and Learning Algorithms for Robotics. arXiv. preprint Ch. 31
- (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 29
- (2021b). Open Problem: Tight Online Confidence Intervals for RKHS Elements. Conference on Learning Theory. Ch. 29
- (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. Ch. 27 Ch. 29
- (2023b). Recent Advances in Bayesian Optimization. ACM Computing Surveys. Ch. 31
- (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. Ch. 27 Ch. 28
- (2025b). Fusing Reward and Dueling Feedback in Stochastic Bandits. International Conference on Machine Learning. Ch. 28
- (2025d). Personalized Building Climate Control with Contextual Preferential Bayesian Optimization. arXiv. preprint Ch. 28
- (2023). On the Sublinear Regret of GP-UCB. Advances in Neural Information Processing Systems. Ch. 29
- (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Ch. 26
- (2024). Stopping Bayesian Optimization with Probabilistic Regret Bounds. NeurIPS 2024. Ch. 30
- (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Ch. 26 Ch. 28 Ch. 29 Ch. 31
- (2024). Borda Regret Minimization for Generalized Linear Dueling Bandits. International Conference on Machine Learning. Ch. 29
- (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. preprint Ch. 27
- (2024). Cost-aware Bayesian Optimization via the Pandora's Box Gittins Index. NeurIPS 2024. Ch. 30
- (2026). Cost-aware Stopping for Bayesian Optimization. International Conference on Machine Learning. Ch. 30
- (2025). Bayesian Optimization with Constraints, Structure and Human Feedback. École Polytechnique Fédérale de Lausanne (EPFL). doi:10.5075/epfl-thesis-11166. thesis Ch. 31
- (2020a). Preference-based Reinforcement Learning with Finite-Time Guarantees. Advances in Neural Information Processing Systems. Ch. 29
- (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. Ch. 28 Ch. 29
- (2024a). Principled Bayesian Optimisation in Collaboration with Human Experts. NeurIPS 2024. Ch. 30
- (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 31
- (2025b). Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization. ICLR 2025 (oral). Ch. 26 Ch. 30
- (2026). GIT-BO: High-Dimensional Bayesian Optimization with Tabular Foundation Models. International Conference on Learning Representations. Ch. 30
- (2026). Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery. ICLR 2026. Ch. 30
- (2009). Interactively optimizing information retrieval systems as a dueling bandits problem. Proceedings of the 26th Annual International Conference on Machine Learning. Ch. 26 Ch. 29
- (2012). The K-armed Dueling Bandits Problem. Journal of Computer and System Sciences. Ch. 29
- (2025). PABBO code repository: evaluation config evaluate.yaml. GitHub. software Ch. 27 Ch. 30 Ch. 31
- (2026). PABBO. GitHub. software Ch. 31
- (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
- (2026a). In-Context Multi-Objective Optimization. International Conference on Learning Representations. Ch. 30
- (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Ch. 27
- (2024). Global and preference-based optimization using surrogate-based methods. IMT School for Advanced Studies Lucca. doi:10.13118/imtlucca/e-theses/415. thesis Ch. 31
- (2025). PWAS. GitHub. software Ch. 31
- (2025). Global and Preference-Based Optimization with Mixed Variables Using Piecewise Affine Surrogates. Journal of Optimization Theory and Applications. Ch. 28
- (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Ch. 27 Ch. 28 Ch. 31
- (2023). Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. International Conference on Machine Learning. Ch. 29
- (2024). Bayesian Optimisation with Unknown Hyperparameters: Regret Bounds Logarithmically Closer to Optimal. NeurIPS 2024. Ch. 30