Bayesian Optimization
Part VI: The Research Frontier
中文

The Research Frontier

The first four parts present the methods as they are usually taught. This part reports what nine years of research since González et al. (2017) have established about them, what remains contested, and what is missing, from a systematic reading of the literature through September 2026. It covers the history of the field, the models behind a human answer, the rules for choosing queries and their documented failures, the theory of learning from comparisons, the scaling of the methods to many dimensions, and the software and evaluation practices that shape every published result.

Read together, the chapters show where the field's difficulty now lies. The algorithms matured over the decade, with a decision-theoretic foundation for choosing queries and regret bounds that caught up with scalar feedback, while the default model of a human answer stayed as it was in 2005. The bottleneck has moved from algorithms to measurement: what a single comparison measures, how answers should be modeled, and what asking does to the person who answers (Section 45.1).

The tone changes accordingly. Claims carry their evidence: venues, sample sizes, the conditions under which a result was observed, and whether it has been peer reviewed. The book's own inferences are marked. Each chapter ends by sorting its conclusions into what is settled, what is contested, and what is missing.

The part assumes Part IV.

Chapters in this part

  1. 26 A Decade of Preferential Bayesian Optimization

    From the 2005 baselines to September 2026: how the field got its name, how its tools and inference settled, the decision-theoretic turn, and the years in which theory caught up and the default pipeline came under scrutiny. An interactive timeline places every milestone in its lane and phase.

  2. 27 Observation Models, Surrogates, and Inference

    What the likelihood assumes about a human answer, which surrogates replace the Gaussian process and why, how much the inference approximation matters, and what the default implementation actually does.

  3. 28 Acquisition, Query Forms, and Problem Extensions

    How the rules for choosing queries evolved from heuristics to decision theory, the failure modes several groups found independently, the forms a query can take and the problem variants built on preferential Bayesian optimization, and why the published comparisons, each run at its own dimension and noise level, cannot simply be pooled.

  4. 29 Theory: From Dueling Bandits to Kernelized Preference Optimization

    What is proved about learning from comparisons: the finite-arm and linear dueling-bandit results, the kernelized regret bounds of 2021 to 2026 with their links, assumptions, and regret units, the decision-theoretic results for EUBO, the missing lower bounds, the theory of the observation model, identifiability, and drift, contamination, response times, and stopping.

  5. 30 High Dimensions and the Changing Landscape of Bayesian Optimization

    Why Bayesian optimization was said to fail beyond 10 to 20 dimensions, what scalar BO learned about lengthscale priors and why, how far local preferential methods reach and what confounds them, and where pretrained surrogates, language models, and cost-aware stopping stand for comparisons.

  6. 31 Software, Evaluation, and the Research Community

    The maintained software for PBO and the defaults it ships, why research code is hard to rerun, how methods are evaluated with simulated users and why the choice of metric decides the winner, what changes when real people answer, and who does this research, in which disciplines, and how much of it there is.

References for Part VI

271 works cited across this part's chapters.

  1. Aalto PML (2022). PPBO. GitHub. software Ch. 31
  2. Abdolshah, M., Shilton, A., Rana, S., Gupta, S., and Venkatesh, S. (2019). Multi-objective Bayesian optimisation with preferences over objectives. Advances in Neural Information Processing Systems. Ch. 28
  3. Abeille, M., Faury, L., and Calauzènes, C. (2021). Instance-Wise Minimax-Optimal Algorithms for Logistic Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 29
  4. Adachi, M., Chau, S. L., Xu, W., Singh, A., Osborne, M. A., and Muandet, K. (2025). Bayesian Optimization for Building Social-Influence-Free Consensus. arXiv. preprint Ch. 28
  5. Agarwal, A., Agarwal, S., and Patil, P. (2021). Stochastic Dueling Bandits with Adversarial Corruption. Algorithmic Learning Theory. Ch. 29
  6. Agarwal, A., Ghuge, R., and Nagarajan, V. (2022). Batched Dueling Bandits. International Conference on Machine Learning. Ch. 29
  7. Agnihotri, A., Jain, R., Ramachandran, D., and Wen, Z. (2026). Best Policy Learning From Trajectory Preference Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 29
  8. An, Z., Nakshbandi, D., and Du, W. (2026). Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences. arXiv. preprint Ch. 29
  9. arXiv (2026a). Abstract search: preference terms AND "Bayesian optimization". arXiv API. non-peer-reviewed Ch. 31
  10. arXiv (2026b). Abstract search: preferential AND Bayesian AND (optimization OR optimisation). arXiv API. non-peer-reviewed Ch. 26 Ch. 31
  11. Astudillo, R. (2023a). qEUBO. GitHub. software Ch. 31
  12. Astudillo, R. (2023b). qEUBO author code repository: noise-level calibration script get_noise_level.py (the calibrated Ackley noise levels are set in experiments/ackley_runner.py). GitHub. software Ch. 28 Ch. 31
  13. Astudillo, R., and Frazier, P. (2020). Multi-attribute Bayesian optimization with interactive preference learning. International Conference on Artificial Intelligence and Statistics. Ch. 28 Ch. 31
  14. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 30 Ch. 31
  15. Astudillo, R., Li, K., Tucker, M., Cheng, C. X., Ames, A. D., and Yue, Y. (2025). Preferential Multi-Objective Bayesian Optimization. Transactions on Machine Learning Research. Ch. 27 Ch. 28 Ch. 31
  16. Astudillo Marban, R. (2022). Exploiting Composite Functions in Bayesian Optimization. Cornell University. thesis Ch. 31
  17. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Ch. 26 Ch. 28
  18. AutoML.org (2026). smac 2.4.1. PyPI. software Ch. 31
  19. Bemporad (2023). GLIS. GitHub. software Ch. 31
  20. Bemporad, A., and Piga, D. (2021). Global optimization based on active preference learning with radial basis functions. Machine Learning. Ch. 26 Ch. 27 Ch. 31
  21. Benavoli, A., and Azzimonti, D. (2024). Linearly Constrained Gaussian Processes are SkewGPs: application to Monotonic Preference Learning and Desirability. Uncertainty in Artificial Intelligence. Ch. 27
  22. Benavoli, A., and Azzimonti, D. (2026a). A tutorial on learning from preferences and choices with Gaussian Processes. Foundations and Trends in Machine Learning 19(1):1-120. Ch. 26 Ch. 27 Ch. 31
  23. Benavoli, and Azzimonti (2026b). prefGP. GitHub. software Ch. 31
  24. Benavoli, A., Azzimonti, D., and Piga, D. (2020). Skew Gaussian processes for classification. Machine Learning. Ch. 27 Ch. 29
  25. Benavoli, A., Azzimonti, D., and Piga, D. (2021a). A unified framework for closed-form nonparametric regression, classification, preference and mixed problems with Skew Gaussian Processes. Machine Learning. Ch. 27 Ch. 29
  26. Benavoli, A., Azzimonti, D., and Piga, D. (2021b). Choice functions based multi-objective Bayesian optimisation. arXiv. preprint Ch. 28
  27. Benavoli, A., Azzimonti, D., and Piga, D. (2021c). Preferential Bayesian optimisation with skew gaussian processes. Proceedings of the Genetic and Evolutionary Computation Conference Companion. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 31
  28. Benavoli, A., Azzimonti, D., and Piga, D. (2023). Learning Choice Functions with Gaussian Processes. Uncertainty in Artificial Intelligence. Ch. 27 Ch. 28
  29. Benavoli, A., Azzimonti, D., and Piga, D. (2025). SkewGP. GitHub. software Ch. 31
  30. Bengs, V., Busa-Fekete, R., El Mesaoudi-Paul, A., and Hüllermeier, E. (2021). Preference-based Online Learning with Dueling Bandits: A Survey. Journal of Machine Learning Research. Ch. 26 Ch. 29 Ch. 31
  31. Bengs, V., Saha, A., and Hüllermeier, E. (2022). Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity Models. International Conference on Machine Learning. Ch. 29
  32. Bengs, V., Haddenhorst, B., and Hüllermeier, E. (2024). Identifying Copeland Winners in Dueling Bandits with Indifferences. International Conference on Artificial Intelligence and Statistics. Ch. 29
  33. Benkert, J.-M., Liu, S., and Netzer, N. (2026). Time is Knowledge: What Response Times Reveal. working paper (arXiv). working paper Ch. 29
  34. Bergna, R., Depeweg, S., and Hernández-Lobato, J. M. (2026). Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors. arXiv. preprint Ch. 30
  35. Bıyık, E., Palan, M., Landolfi, N. C., Losey, D. P., and Sadigh, D. (2019). Asking Easy Questions: A User-Friendly Approach to Active Reward Learning. CoRL 2019. Ch. 26 Ch. 27 Ch. 28 Ch. 30
  36. Bıyık, E., Huynh, N., Kochenderfer, M. J., and Sadigh, D. (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. RSS 2020. Ch. 27
  37. Blum, A., Gupta, M., Li, G., Manoj, N. S., Saha, A., and Yang, Y. (2024). Dueling Optimization with a Monotone Adversary. International Conference on Algorithmic Learning Theory. Ch. 29
  38. Bogunovic, I., Scarlett, J., and Cevher, V. (2016). Time-Varying Gaussian Process Bandit Optimization. AISTATS 2016. Ch. 29
  39. Bogunovic, I., Krause, A., and Scarlett, J. (2020). Corruption-Tolerant Gaussian Process Bandit Optimization. International Conference on Artificial Intelligence and Statistics. Ch. 29
  40. Brochu, E., de Freitas, N., and Ghosh, A. (2007). Active Preference Learning with Discrete Choice Data. Advances in Neural Information Processing Systems. Ch. 26 Ch. 27 Ch. 28
  41. Brochu, E., Cora, V. M., and de Freitas, N. (2010). A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning. arXiv preprint. preprint Ch. 31
  42. Bukharin, A., Hong, I., Jiang, H., Li, Z., Zhang, Q., Zhang, Z., and Zhao, T. (2024). Robust Reinforcement Learning from Corrupted Human Feedback. Advances in Neural Information Processing Systems. Ch. 29
  43. Cai, X., and Scarlett, J. (2021). On Lower Bounds for Standard and Robust Gaussian Process Bandit Optimization. International Conference on Machine Learning. Ch. 29
  44. Cao, L., Shi, M., and Shroff, N. B. (2026). Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries. UAI 2026. Ch. 29
  45. Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. Ch. 26
  46. Chan, L., Liao, Y.-C., Mo, G. B., Dudley, J. J., Cheng, C.-L., Kristensson, P. O., and Oulasvirta, A. (2022). Investigating Positive and Negative Qualities of Human-in-the-Loop Optimization for Designing Interaction Techniques. CHI 2022. Ch. 26 Ch. 31
  47. Chau, S. L., González, J., and Sejdinovic, D. (2022). Learning Inconsistent Preferences with Gaussian Processes. International Conference on Artificial Intelligence and Statistics. Ch. 27 Ch. 29
  48. Chen, B., and Frazier, P. I. (2017). Dueling Bandits with Weak Regret. International Conference on Machine Learning. Ch. 29
  49. Chen, X., Zhong, H., Yang, Z., Wang, Z., and Wang, L. (2022). Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation. International Conference on Machine Learning. Ch. 29
  50. Chen, E., Truong, S. T., Dullerud, N., Koyejo, S., and Guestrin, C. (2026). Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds. Conference on Uncertainty in Artificial Intelligence. Ch. 28
  51. Cheng, M., Novoseller, E., Tucker, M., Cheng, R., Yue, Y., and Burdick, J. (2020). Preference-Based Bayesian Optimization in High Dimensions with Human Feedback. SCMLS 2020 Workshop. workshop paper Ch. 28
  52. Chidambaram, K., Seetharaman, K. V., and Syrgkanis, V. (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Ch. 29
  53. Chowdhury, S. R., and Gopalan, A. (2017). On Kernelized Multi-armed Bandits. International Conference on Machine Learning. Ch. 29
  54. Chu, W., and Ghahramani, Z. (2005). Preference learning with Gaussian processes. Proceedings of the 22nd international conference on Machine learning - ICML '05. Ch. 26 Ch. 27
  55. Colella, F., Daee, P., Jokinen, J., Oulasvirta, A., and Kaski, S. (2020). Human Strategic Steering Improves Performance of Interactive Optimization. UMAP 2020. Ch. 31
  56. Cosner, R., Tucker, M., Taylor, A., Li, K., Molnár, T., Ubelacker, W., … Ames, A. (2022). Safety-Aware Preference-Based Learning for Safety-Critical Control. Learning for Dynamics and Control Conference. Ch. 28
  57. Coutinho, J. P. L., Peng, Y., Rendall, R., Rizzo, C., Ma, K., Chin, S.-T., Castillo, I., and Reis, M. S. (2025). Accelerated controller tuning using human feedback and Multi-Task Preferential Bayesian Optimization. 2025 American Control Conference (ACC). Ch. 28
  58. Coutinho, J. P., Peng, Y., Rendall, R., Ma, K., Chin, S.-T., Castillo, I., and Reis, M. S. (2026). Efficient human-in-the-loop MPC tuning with multi-task preferential Bayesian optimization. Control Engineering Practice. Ch. 28
  59. CyberAgent AI Lab (2023). preferentialBO. GitHub. software Ch. 31
  60. Dao, L. A., Maccarini, M., Nicora, M. L., Falerni, M. M., Mondellini, M., Veerappan, P., … Roveda, L. (2025). Experience in Engineering Complex Systems: Active Preference Learning With Multiple Outcomes and Certainty Levels. IEEE Transactions on Human-Machine Systems. Ch. 27
  61. De Peuter, S., Zhu, S., Guo, Y., Howes, A., and Kaski, S. (2024). Preference Learning of Latent Decision Utilities with a Human-like Model of Preferential Choice. Advances in Neural Information Processing Systems. Ch. 29
  62. Di, Q., Jin, T., Wu, Y., Zhao, H., Farnoud, F., and Gu, Q. (2024). Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits. International Conference on Learning Representations. Ch. 29
  63. Di, Q., He, J., and Gu, Q. (2025). Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback. International Conference on Machine Learning. Ch. 29
  64. Doumont, C., Fan, D., Maus, N., Gardner, J. R., Moss, H., and Pleiss, G. (2026). We Still Don't Understand High-Dimensional Bayesian Optimization. AISTATS 2026 (best student paper). Ch. 26 Ch. 30
  65. Drago, S., Mussi, M., and Metelli, A. M. (2025). Towards Theoretical Understanding of Sequential Decision Making with Preference Feedback. International Conference on Machine Learning. Ch. 29
  66. Dragonfly developers (2022). dragonfly-opt 0.1.7. PyPI. software Ch. 31
  67. Dubey, M., De Peuter, S., Wang, W., and Kaski, S. (2026). Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization. arXiv. preprint Ch. 27
  68. Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M. (2015). Contextual Dueling Bandits. Conference on Learning Theory. Ch. 29
  69. Durante, D. (2019). Conjugate Bayes for probit regression via unified skew-normal distributions. Biometrika. Ch. 29
  70. Dwaracherla, V., Asghari, S. M., Hao, B., and Van Roy, B. (2024). Efficient Exploration for LLMs. ICML 2024. Ch. 26
  71. Emukit developers (2026). preferential_batch_bayesian_optimization example. GitHub. software Ch. 31
  72. Erarslan, A., Sevilla Salcedo, C., Tanskanen, V., Nisov, A., Päiväkumpu, E., Aisala, H., … Mikkola, P. (2025). Consecutive Preferential Bayesian Optimization. arXiv. preprint Ch. 27 Ch. 28
  73. Facebook, Inc. (2022). ax-platform 0.2.6. PyPI. software Ch. 26 Ch. 31
  74. Fan, D., and Pleiss, G. (2026). Adaptive Candidate Point Thompson Sampling for High-Dimensional Bayesian Optimization. AISTATS 2026. Ch. 30
  75. Faury, L., Abeille, M., Calauzènes, C., and Fercoq, O. (2020). Improved Optimistic Algorithms for Logistic Bandits. International Conference on Machine Learning. Ch. 29
  76. Fauvel, T. (2021). Human-in-the-loop optimization of retinal prostheses encoders. Sorbonne Université. thesis Ch. 31
  77. Fauvel, T., and Chalk, M. (2021). Efficient Exploration in Binary and Preferential Bayesian Optimization. arXiv. preprint Ch. 27 Ch. 28 Ch. 31
  78. FiveThirtyEight (2017). candy-power-ranking data. GitHub. non-peer-reviewed Ch. 31
  79. Frazier, P. I. (2018). A Tutorial on Bayesian Optimization. arXiv. preprint Ch. 30 Ch. 31
  80. Gardner, J. R., Kusner, M. J., Xu, Z., Weinberger, K. Q., and Cunningham, J. P. (2014). Bayesian Optimization with Inequality Constraints. Proceedings of the 31st International Conference on Machine Learning (ICML 2014). Ch. 28
  81. Garnett, R. (2023). Bayesian Optimization. Cambridge University Press. Ch. 31
  82. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 31
  83. GPflow developers (2026). gpflow 2.11.1. PyPI. software Ch. 31
  84. GPyTorch developers (2026). gpytorch 1.15.2. PyPI. software Ch. 31
  85. Granley, J., Fauvel, T., Chalk, M., and Beyeler, M. (2023). Human-in-the-Loop Optimization for Deep Stimulus Encoding in Visual Prostheses. NeurIPS 2023. Ch. 27 Ch. 28 Ch. 30 Ch. 31
  86. Gupta, R., Hartford, J., and Liu, B. (2025). LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Findings of the Association for Computational Linguistics: EMNLP 2025. Ch. 30
  87. Haddenhorst, B., Bengs, V., and Hüllermeier, E. (2021a). Identification of the Generalized Condorcet Winner in Multi-dueling Bandits. Advances in Neural Information Processing Systems. Ch. 29
  88. Haddenhorst, B., Bengs, V., Brandt, J., and Hüllermeier, E. (2021b). Testification of Condorcet Winners in dueling bandits. Uncertainty in Artificial Intelligence. Ch. 29
  89. Haltia, A., Hyvönen, V., and Kaski, S. (2026). Elicitation-Augmented Bayesian Optimization. arXiv. preprint Ch. 28
  90. Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M. (2011). Bayesian Active Learning for Classification and Preference Learning. arXiv. preprint Ch. 27 Ch. 28
  91. Houlsby, N., Huszár, F., Ghahramani, Z., and Hernández-lobato, J. (2012). Collaborative Gaussian Processes for Preference Learning. Advances in Neural Information Processing Systems. Ch. 27
  92. Huawei Noah's Ark Lab (2024). HEBO 0.3.6. PyPI. software Ch. 31
  93. Huber, F., Rojas Gonzalez, S., and Astudillo, R. (2025). Bayesian Preference Elicitation for Decision Support in Multi‐Objective Optimization. Journal of Multi-Criteria Decision Analysis. Ch. 28
  94. Hvarfner, C., Hellsten, E. O., and Nardi, L. (2024). Vanilla Bayesian Optimization Performs Great in High Dimensions. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 30
  95. Hvarfner, C., Eriksson, D., Bakshy, E., and Balandat, M. (2025). Informed Initialization for Bayesian Optimization and Active Learning. NeurIPS 2025. Ch. 30
  96. Hvarfner, C., Daulton, S., Balandat, M., and Bakshy, E. (2026). Pitfalls and Remedies for Multi-Task Bayesian Optimization. arXiv. preprint Ch. 30
  97. ICML (2023). The Many Facets of Preference-Based Learning. ICML 2023 workshop page. non-peer-reviewed Ch. 26 Ch. 31
  98. Ignatenko, T., Kondrashov, K., Cox, M., and de Vries, B. (2025). On preference learning based on sequential Bayesian optimization with pairwise comparison. Artificial Intelligence. Ch. 28
  99. Institute for Data Science in Mechanical Engineering, RWTH Aachen University (2026). crashpbo. GitHub. software Ch. 31
  100. Ip, J. H. S., Chakrabarty, A., Mesbah, A., and Romeres, D. (2025). User Preference Meets Pareto-Optimality in Multi-Objective Bayesian Optimization. Proceedings of the AAAI Conference on Artificial Intelligence. Ch. 28
  101. Ishibashi, H., Karasuyama, M., Takeuchi, I., and Hino, H. (2023). A stopping criterion for Bayesian optimization by the gap of expected minimum simple regrets. International Conference on Artificial Intelligence and Statistics. Ch. 30
  102. Iwai, K., Kumagae, Y., Koyama, Y., Hamasaki, M., and Goto, M. (2025). Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. Ch. 28 Ch. 31
  103. Iwazaki, S., and Takeno, S. (2025). Near-Optimal Algorithm for Non-Stationary Kernelized Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 29
  104. Kamishima, T. (2026). SUSHI Preference Data Sets. kamishima.net. non-peer-reviewed Ch. 31
  105. Kayal (2025). BOHF_code_submission. GitHub. software Ch. 31
  106. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 30 Ch. 31
  107. Khan, F. A., Chakraborty, T., Dietrich, J. P., and Wirth, C. (2025). Efficient Contextual Preferential Bayesian Optimization with Historical Examples. Proceedings of the Genetic and Evolutionary Computation Conference Companion. Ch. 28
  108. Kirschner, J., and Krause, A. (2021). Bias-Robust Bayesian Optimization via Dueling Bandits. International Conference on Machine Learning. Ch. 26 Ch. 28 Ch. 29 Ch. 31
  109. Kleine Buening, T., and Saha, A. (2023). ANACONDA: An Improved Dynamic Regret Algorithm for Adaptive Non-Stationary Dueling Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 29
  110. Kobalczyk, K., Lin, Z. J., Letham, B., Zhao, Z., Balandat, M., and Bakshy, E. (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Ch. 26 Ch. 28 Ch. 30 Ch. 31
  111. Kolpaczki, P., Bengs, V., and Hüllermeier, E. (2022). Non-Stationary Dueling Bandits. arXiv. preprint Ch. 29
  112. Komiyama, J., Honda, J., Kashima, H., and Nakagawa, H. (2015). Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem. Conference on Learning Theory. Ch. 29
  113. Koyama, Y. (2017). Computational Design Driven by Visual Aesthetic Preference. The University of Tokyo. doi:10.15083/00076184. thesis Ch. 31
  114. Koyama, Y. (2025a). preference-regressor.hpp. GitHub. software Ch. 31
  115. Koyama, Y. (2025b). sequential-line-search. GitHub. software Ch. 31
  116. Koyama, Y., and Igarashi, T. (2018). Computational Design with Crowds. Computational Interaction. Ch. 31
  117. Koyama, Y., Sato, I., Sakamoto, D., and Igarashi, T. (2017). Sequential line search for efficient visual design optimization by crowds. ACM Transactions on Graphics. Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
  118. Koyama, Y., Sato, I., and Goto, M. (2020). Sequential Gallery for Interactive Visual Design Optimization. ACM Transactions on Graphics 39(4) (SIGGRAPH 2020). Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
  119. Kumagai, W. (2017). Regret Analysis for Continuous Dueling Bandit. Advances in Neural Information Processing Systems. Ch. 26 Ch. 29
  120. Kuss, M., and Rasmussen, C. E. (2005). Assessing Approximate Inference for Binary Gaussian Process Classification. Journal of Machine Learning Research. Ch. 27
  121. Kwon, Y., Tsurumine, Y., Shimmura, T., Kawamura, S., and Matsubara, T. (2022). Physically Consistent Preferential Bayesian Optimization for Food Arrangement. IEEE Robotics and Automation Letters. Ch. 28
  122. Landolt, L., Maddux, A. M., Schlaginhaufen, A., Vaishampayan, S., and Kamgarpour, M. (2026). Eliciting Truthful Feedback for Preference-Based Learning via the VCG Mechanism. International Conference on Artificial Intelligence and Statistics. Ch. 29
  123. Langerak, T., Zhang, R., Wang, Z., Kristensson, P. O., and Oulasvirta, A. (2026). Cost-Aware Bayesian Optimization for Prototyping Interactive Devices. CHI 2026. Ch. 26
  124. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 26 Ch. 28 Ch. 29 Ch. 31
  125. Lee, J., Yi, S.-w., and Oh, M.-h. (2025a). Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options. NeurIPS 2025. Ch. 29
  126. Leenders, N., Quadt, T., Cule, B., Lindelauf, R., Monsuur, H., van Oijen, J., and Voskuijl, M. (2025). DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization. arXiv. preprint Ch. 27
  127. Li, Z., and Scarlett, J. (2022). Gaussian Process Bandit Optimization with Few Batches. International Conference on Artificial Intelligence and Statistics. Ch. 29
  128. Li, K., Tucker, M., Bıyık, E., Novoseller, E., Burdick, J. W., Sui, Y., … Ames, A. D. (2021). ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes. ICRA 2021. Ch. 27 Ch. 28
  129. Li, W., Rinaldo, A., and Wang, D. (2022). Detecting Abrupt Changes in Sequential Pairwise Comparison Data. Advances in Neural Information Processing Systems. Ch. 29
  130. Li, S., Zhang, Y., Ren, Z., Liang, C., Li, N., and Shah, J. A. (2024a). Enhancing Preference-based Linear Bandits via Human Response Time. Advances in Neural Information Processing Systems. Ch. 27 Ch. 29
  131. Li, X., Zhao, H., and Gu, Q. (2024b). Feel-Good Thompson Sampling for Contextual Dueling Bandits. International Conference on Machine Learning. Ch. 29
  132. Li, Z., Liao, Y.-C., and Holz, C. (2025a). Efficient Visual Appearance Optimization by Learning from Prior Preferences. UIST 2025. Ch. 26 Ch. 27 Ch. 28
  133. Liao, Y.-C., Belo, J., Moon, H.-S., Steimle, J., and Feit, A. M. (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Ch. 26
  134. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Ch. 26 Ch. 28 Ch. 31
  135. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Ch. 27 Ch. 28 Ch. 30
  136. Liu, M., Chen, Y., Fan, Z., Farina, G., Ozdaglar, A., and Zhang, K. (2026c). Online Learning and Equilibrium Computation with Ranking Feedback. ICLR 2026. Ch. 29
  137. Liu, K., Long, Q., Shi, Z., Su, W. J., and Xiao, J. (2026e). Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium. The Annals of Statistics. doi:10.1214/26-aos2643. Ch. 29
  138. Mandal, D., Nika, A., Kamalaruban, P., Singla, A., and Radanovic, G. (2025). Corruption Robust Offline Reinforcement Learning with Human Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 29
  139. Maran, D., Bacchiocchi, F., Stradi, F. E., Castiglioni, M., Gatti, N., and Restelli, M. (2024). Bandits with Ranking Feedback. Advances in Neural Information Processing Systems. Ch. 29
  140. McCourt, M., and Dewancker, I. (2019). Sampling Humans for Optimizing Preferences in Coloring Artwork. ICML 2019 Workshop on Human in the Loop Learning. workshop paper Ch. 31
  141. Meindl, J., Tian, Y., Cui, T., Thost, V., Hong, Z.-W., Dürholt, J., … Luković, M. K. (2025). ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization. arXiv. preprint Ch. 30
  142. Menn, J., Kober, M., Brunzema, P., Stenger, D., and Trimpe, S. (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
  143. Menn, J., Stenger, D., and Trimpe, S. (2026b). Preferential Bayesian Optimization with Crash Feedback. IEEE Robotics and Automation Letters. doi:10.1109/LRA.2026.3665446. Ch. 26 Ch. 27 Ch. 28 Ch. 31
  144. Meta (2026). AEPsych. GitHub. software Ch. 31
  145. Meta Platforms, Inc. (2026a). ax-platform release history. PyPI. software Ch. 31
  146. Meta Platforms, Inc. (2026b). ax/generation_strategy/transition_criterion.py. GitHub. software Ch. 31
  147. Meta Platforms, Inc. (2026c). Bayesian optimization with pairwise comparison data (preferential Bayesian optimization tutorial, documentation v0.18.1). botorch.org. software Ch. 28 Ch. 31
  148. Meta Platforms, Inc. (2026d). Bayesian optimization with preference exploration (BOPE tutorial, documentation v0.18.1). botorch.org. software Ch. 28
  149. Meta Platforms, Inc. (2026e). BoTorch CHANGELOG. GitHub. software Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
  150. Meta Platforms, Inc. (2026f). BoTorch LICENSE. GitHub. software Ch. 31
  151. Meta Platforms, Inc. (2026g). BoTorch pairwise likelihood source code likelihoods/pairwise.py. GitHub. software Ch. 27
  152. Meta Platforms, Inc. (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Ch. 27 Ch. 30 Ch. 31
  153. Meta Platforms, Inc. (2026i). botorch release history. PyPI. software Ch. 31
  154. Meta Platforms, Inc. (2026j). botorch/acquisition/preference.py. GitHub. software Ch. 31
  155. Meta Platforms, Inc. (2026k). botorch/models/utils/gpytorch_modules.py. GitHub. software Ch. 30 Ch. 31
  156. Meta Platforms, Inc. (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Ch. 26 Ch. 31
  157. Meta Platforms, Inc. (2026m). tutorials directory. GitHub. software Ch. 31
  158. Meta Research (2023). qEUBO. GitHub. software Ch. 31
  159. Meta Research (2026). lilo. GitHub. software Ch. 31
  160. Mikkola, P. (2024). Humans as Information Sources in Bayesian Optimization. Aalto University. thesis Ch. 31
  161. Mikkola, P., Todorović, M., Järvi, J., Rinke, P., and Kaski, S. (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
  162. Müller, S., Reuter, A., Hollmann, N., Rügamer, D., and Hutter, F. (2025). Position: The Future of Bayesian Prediction Is Prior-Fitted. ICML 2025 (position paper). Ch. 30
  163. Nguyen, Q. P., Tay, S., Low, B. K. H., and Jaillet, P. (2021). Top- Ranking Bayesian Optimization. AAAI 2021. Ch. 27 Ch. 28
  164. Niwa, R., Yoshida, S., Koyama, Y., and Ushiku, Y. (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Ch. 26
  165. Novoseller, E. R. (2021). Online Learning from Human Feedback with Applications to Exoskeleton Gait Optimization. California Institute of Technology. doi:10.7907/gvtx-1586. thesis Ch. 31
  166. Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J. (2020). Dueling Posterior Sampling for Preference-Based Reinforcement Learning. Conference on Uncertainty in Artificial Intelligence. Ch. 29
  167. Oh, Y. (2026). Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions. ICML 2026. Ch. 29
  168. Oh, Y., Park, J., and Paik, T. (2026a). Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration. International Conference on Artificial Intelligence and Statistics. Ch. 29
  169. Olson, M., Santorella, E., Tiao, L. C., Cakmak, S., Garrard, M., Daulton, S., … Bakshy, E. (2025). Ax: A Platform for Adaptive Experimentation. International Conference on Automated Machine Learning. Ch. 31
  170. Optuna developers (2026a). optuna 5.0.0. PyPI. software Ch. 31
  171. Optuna developers (2026b). optuna-dashboard 0.21.0. PyPI. software Ch. 26 Ch. 30 Ch. 31
  172. Optuna developers (2026c). optuna-dashboard PreferentialGPSampler source code gp.py. GitHub. software Ch. 27 Ch. 30 Ch. 31
  173. Ou, C., Buschek, D., Mayer, S., and Butz, A. (2022). The Human in the Infinite Loop: A Case Study on Revealing and Explaining Human-AI Interaction Loop Failures. Mensch und Computer 2022. Ch. 26 Ch. 31
  174. Ou, C., Mayer, S., and Butz, A. (2023). The Impact of Expertise in the Loop for Exploring Machine Rationality. IUI 2023. Ch. 26 Ch. 30 Ch. 31
  175. Ozaki, R. (2026). PLMBO (Preference Learning Multi-Objective Bayesian Optimization). OptunaHub. software Ch. 31
  176. Ozaki, R., Ishikawa, K., Kanzaki, Y., Takeno, S., Takeuchi, I., and Karasuyama, M. (2024). Multi-Objective Bayesian Optimization with Active Preference Learning. Proceedings of the AAAI Conference on Artificial Intelligence. Ch. 28 Ch. 31
  177. Papenmeier, L., Cheng, N., Becker, S., and Nardi, L. (2025a). Exploring Exploration in Bayesian Optimization. Conference on Uncertainty in Artificial Intelligence. Ch. 30
  178. Papenmeier, L., Poloczek, M., and Nardi, L. (2025b). Understanding High-Dimensional Bayesian Optimization. ICML 2025, PMLR 267:47902-47923. Ch. 26 Ch. 30
  179. Pásztor, B., Kassraie, P., and Krause, A. (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Ch. 26 Ch. 28 Ch. 29 Ch. 31
  180. Peng, S., Chen, H., and Driggs-Campbell, K. (2025). Towards Uncertainty Unification: A Case Study for Preference Learning. RSS 2025. Ch. 27
  181. Pith (2026). Machine-generated review of arXiv 2505.23673 (MR-LPF). pith.science. non-peer-reviewed Ch. 29
  182. PREDICT-EPFL (2024). POP-BO. GitHub. software Ch. 31
  183. Previtali, D., Mazzoleni, M., Ferramosca, A., and Previdi, F. (2023). GLISp-r: a preference-based optimization algorithm with convergence guarantees. Computational Optimization and Applications. Ch. 27
  184. ProbML (2026). Symposium on Probabilistic Machine Learning website. probml.cc. non-peer-reviewed Ch. 31
  185. Pukdee, R., Balcan, M.-F., and Ravikumar, P. (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. Ch. 27
  186. Quadt, T. (2026). DT-PBO-preprint. GitHub. software Ch. 31
  187. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. Ch. 26
  188. Ranković, B., Griffiths, R.-R., and Schwaller, P. (2026). Large language models as uncertainty-calibrated optimizers for experimental discovery. Nature Machine Intelligence. doi:10.1038/s42256-026-01283-z. Ch. 30
  189. Rogers, T., and Ponnada, S. (2026). Zero-shot Bayesian optimization with TabPFN: Competitive with state-of-the-art without per-task training. AutoML Conference 2026 (per Amazon Science page). Ch. 30
  190. Saad, E. M., Carpentier, A., Kocák, T., and Verzelen, N. (2024). On Weak Regret Analysis for Dueling Bandits. Advances in Neural Information Processing Systems. Ch. 29
  191. Saha, A. (2021). Optimal Algorithms for Stochastic Contextual Preference Bandits. Advances in Neural Information Processing Systems. Ch. 29
  192. Saha, A., and Gaillard, P. (2021). Dueling Bandits with Adversarial Sleeping. Advances in Neural Information Processing Systems. Ch. 29
  193. Saha, A., and Gaillard, P. (2022). Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences. International Conference on Machine Learning. Ch. 29
  194. Saha, A., and Gopalan, A. (2019a). Combinatorial Bandits with Relative Feedback. Advances in Neural Information Processing Systems. Ch. 29
  195. Saha, A., and Gopalan, A. (2019b). PAC Battling Bandits in the Plackett-Luce Model. Algorithmic Learning Theory. Ch. 29
  196. Saha, A., and Gopalan, A. (2020). From PAC to Instance-Optimal Sample Complexity in the Plackett-Luce Model. International Conference on Machine Learning. Ch. 29
  197. Saha, A., and Gupta, S. (2022). Optimal and Efficient Dynamic Regret Algorithms for Non-Stationary Dueling Bandits. International Conference on Machine Learning. Ch. 29
  198. Saha, A., and Krishnamurthy, A. (2022). Efficient and Optimal Algorithms for Contextual Dueling Bandits under Realizability. International Conference on Algorithmic Learning Theory. Ch. 29
  199. Saha, A., Koren, T., and Mansour, Y. (2021a). Adversarial Dueling Bandits. International Conference on Machine Learning. Ch. 29
  200. Saha, A., Koren, T., and Mansour, Y. (2021b). Dueling Convex Optimization. International Conference on Machine Learning. Ch. 29
  201. Saha, A., Feldman, V., Mansour, Y., and Koren, T. (2024). Faster Convergence with MultiWay Preferences. International Conference on Artificial Intelligence and Statistics. Ch. 29
  202. Saha, A., Koren, T., and Mansour, Y. (2025). Dueling Convex Optimization with General Preferences. International Conference on Machine Learning. Ch. 29
  203. Salgia, S., Vakili, S., and Zhao, Q. (2021). A Domain-Shrinking based Bayesian Optimization Algorithm with Order-Optimal Regret Performance. Advances in Neural Information Processing Systems. Ch. 29
  204. Scarlett, J., Bogunovic, I., and Cevher, V. (2017). Lower Bounds on Regret for Noisy Gaussian Process Bandit Optimization. Conference on Learning Theory. Ch. 29
  205. Schäfer, N., Zhao, G., Li, B., Kupnik, M., Seyfarth, A., Beckerle, P., and Grimmer, M. (2026). User preference-based human-in-the-loop tuning of exoskeleton assistance during walking. npj Biomedical Innovations. doi:10.1038/s44385-026-00085-7. Ch. 26
  206. Schoinas, E., Rastogi, A., Carter, A., Granley, J., and Beyeler, M. (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Ch. 31
  207. Secondmind Labs (2026). trieste 4.6.0. PyPI. software Ch. 31
  208. Sekhari, A., Sridharan, K., Sun, W., and Wu, R. (2023). Contextual Bandits and Imitation Learning with Preference-Based Active Queries. Advances in Neural Information Processing Systems. Ch. 29
  209. Semantic Scholar (2026). Bulk search: "preferential bayesian optimization". Semantic Scholar API. non-peer-reviewed Ch. 31
  210. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Ch. 26 Ch. 27 Ch. 28 Ch. 31
  211. SheffieldML (2023). GPyOpt (archived). GitHub. software Ch. 31
  212. Shukla, A., and Basu, D. (2024). Preference-based Pure Exploration. Advances in Neural Information Processing Systems. Ch. 29
  213. Shvartsman, M., Letham, B., Bakshy, E., and Keeley, S. (2024). Response Time Improves Gaussian Process Models for Perception and Preferences. Uncertainty in Artificial Intelligence. Ch. 27 Ch. 29
  214. Siivola, E. (2021). Applications of human feedback in Gaussian processes. Aalto University. thesis Ch. 31
  215. Siivola, E., Dhaka, A. K., Andersen, M. R., González, J., García Moreno, P., and Vehtari, A. (2021). Preferential Batch Bayesian Optimization. IEEE MLSP 2021. Ch. 27 Ch. 28 Ch. 31
  216. Simpson, E., and Gurevych, I. (2020). Scalable Bayesian preference learning for crowds. Machine Learning. Ch. 27
  217. Sinaga, M. A., Martinelli, J., and Kaski, S. (2026). Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization. Symposium on Probabilistic Machine Learning (ProbML 2026), Proceedings Track. Ch. 27 Ch. 28 Ch. 31
  218. Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Ch. 29
  219. Son, S., Bankes, W., Chowdhury, S. R., Paige, B., and Bogunovic, I. (2025). Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift. International Conference on Machine Learning. Ch. 29
  220. Sui, Y., Zhuang, V., Burdick, J. W., and Yue, Y. (2017b). Multi-dueling Bandits with Dependent Arms. UAI 2017. Ch. 26 Ch. 28 Ch. 29
  221. Sui, Y., Zoghi, M., Hofmann, K., and Yue, Y. (2018a). Advancements in Dueling Bandits. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence. doi:10.24963/ijcai.2018/776. Ch. 26 Ch. 29
  222. Sui, Y., Zhuang, V., Burdick, J., and Yue, Y. (2018b). Stagewise Safe Bayesian Optimization with Gaussian Processes. International Conference on Machine Learning. Ch. 26 Ch. 28
  223. Suk, J., and Agarwal, A. (2023). When Can We Track Significant Preference Shifts in Dueling Bandits? Advances in Neural Information Processing Systems. Ch. 29
  224. Taddei, S., Koppen, W., Alfio, E., Nuzzo, S., Flynn, L., Diaz, M. A., … Verstraten, T. (2026). Bayesian Preference Elicitation: Human-In-The-Loop Optimization of An Active Prosthesis. arXiv. preprint Ch. 31
  225. Takeno, S., Nomura, M., and Karasuyama, M. (2022). Preferential Bayesian Optimization with Hallucination Believer. NeurIPS 2022 Workshop on Gaussian Processes, Spatiotemporal Modeling, and Decision-making Systems. workshop paper Ch. 31
  226. Takeno, S., Nomura, M., and Karasuyama, M. (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 31
  227. Tang, M., Zhou, Y., and Huang, C. (2025). Tackling Biased Evaluators in Dueling Bandits. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-2520. Ch. 29
  228. Tatsukawa, Y., Shen, I.-C., Dogan, M. D., Qi, A., Koyama, Y., Shamir, A., and Igarashi, T. (2025). FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization. CHI 2025. Ch. 27
  229. Theiner, L., Hirt, S., Steinke, A., and Findeisen, R. (2025). Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles. ECC 2025. Ch. 31
  230. Theiner, L., Pfefferkorn, M., Zhao, Y., Hirt, S., and Findeisen, R. (2026). Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models. European Control Conference. Ch. 28 Ch. 31
  231. Tucker, M. (2023). Enabling Robust and User-Customized Bipedal Locomotion on Lower-Body Assistive Devices via Hybrid System Theory and Preference-Based Learning. California Institute of Technology. doi:10.7907/j9hk-xa17. thesis Ch. 31
  232. Tucker, M. (2024). POLAR. GitHub. software Ch. 31
  233. Tucker, M., Cheng, M., Novoseller, E., Cheng, R., Yue, Y., Burdick, J. W., and Ames, A. D. (2020a). Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits. IROS 2020. Ch. 26 Ch. 28 Ch. 30 Ch. 31
  234. Tucker, M., Novoseller, E., Kann, C., Sui, Y., Yue, Y., Burdick, J. W., and Ames, A. D. (2020b). Preference-Based Learning for Exoskeleton Gait Optimization. 2020 IEEE International Conference on Robotics and Automation (ICRA). Ch. 26 Ch. 28 Ch. 31
  235. Tucker, M., Li, K., Yue, Y., and Ames, A. D. (2022). POLAR: Preference Optimization and Learning Algorithms for Robotics. arXiv. preprint Ch. 31
  236. Vakili, S., Khezeli, K., and Picheny, V. (2021a). On Information Gain and Regret Bounds in Gaussian Process Bandits. International Conference on Artificial Intelligence and Statistics. Ch. 29
  237. Vakili, S., Scarlett, J., and Javidi, T. (2021b). Open Problem: Tight Online Confidence Intervals for RKHS Elements. Conference on Learning Theory. Ch. 29
  238. Verma, A., Dai, Z., Lin, X., Jaillet, P., and Low, B. K. H. (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. Ch. 27 Ch. 29
  239. Wang, X., Jin, Y., Schmitt, S., and Olhofer, M. (2023b). Recent Advances in Bayesian Optimization. ACM Computing Surveys. Ch. 31
  240. Wang, H., Branke, J., and Poloczek, M. (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. Ch. 27 Ch. 28
  241. Wang, X., Zeng, Q., Zuo, J., Liu, X., Hajiesmaili, M., Lui, J. C., and Wierman, A. (2025b). Fusing Reward and Dueling Feedback in Stochastic Bandits. International Conference on Machine Learning. Ch. 28
  242. Wang, W., Shi, J., and Jones, C. N. (2025d). Personalized Building Climate Control with Contextual Preferential Bayesian Optimization. arXiv. preprint Ch. 28
  243. Whitehouse, J., Ramdas, A., and Wu, S. (2023). On the Sublinear Regret of GP-UCB. Advances in Neural Information Processing Systems. Ch. 29
  244. Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., and Dragan, A. (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Ch. 26
  245. Wilson, J. T. (2024). Stopping Bayesian Optimization with Probabilistic Regret Bounds. NeurIPS 2024. Ch. 30
  246. Wu, K., and Gardner, J. R. (2026). Knowledge Gradient for Preference Learning. arXiv. preprint Ch. 26 Ch. 28 Ch. 29 Ch. 31
  247. Wu, Y., Jin, T., Di, Q., Lou, H., Farnoud, F., and Gu, Q. (2024). Borda Regret Minimization for Generalized Linear Dueling Bandits. International Conference on Machine Learning. Ch. 29
  248. Wu, K., Sanders, C., Letham, B., and Guan, P. (2025a). Mixed Likelihood Variational Gaussian Processes. arXiv. preprint Ch. 27
  249. Xie, Q., Astudillo, R., Frazier, P. I., Scully, Z., and Terenin, A. (2024). Cost-aware Bayesian Optimization via the Pandora's Box Gittins Index. NeurIPS 2024. Ch. 30
  250. Xie, Q., Cai, L., Terenin, A., Frazier, P. I., and Scully, Z. (2026). Cost-aware Stopping for Bayesian Optimization. International Conference on Machine Learning. Ch. 30
  251. Xu, W. (2025). Bayesian Optimization with Constraints, Structure and Human Feedback. École Polytechnique Fédérale de Lausanne (EPFL). doi:10.5075/epfl-thesis-11166. thesis Ch. 31
  252. Xu, Y., Wang, R., Yang, L., Singh, A., and Dubrawski, A. (2020a). Preference-based Reinforcement Learning with Finite-Time Guarantees. Advances in Neural Information Processing Systems. Ch. 29
  253. Xu, Y., Joshi, A., Singh, A., and Dubrawski, A. (2020b). Zeroth Order Non-convex optimization with Dueling-Choice Bandits. Conference on Uncertainty in Artificial Intelligence. Ch. 28 Ch. 29
  254. Xu, W., Adachi, M., Jones, C. N., and Osborne, M. A. (2024a). Principled Bayesian Optimisation in Collaboration with Human Experts. NeurIPS 2024. Ch. 30
  255. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 26 Ch. 27 Ch. 28 Ch. 29 Ch. 31
  256. Xu, Z., Wang, H., Phillips, J. M., and Zhe, S. (2025b). Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization. ICLR 2025 (oral). Ch. 26 Ch. 30
  257. Yu, R. T.-Y., Picard, C., and Ahmed, F. (2026). GIT-BO: High-Dimensional Bayesian Optimization with Tabular Foundation Models. International Conference on Learning Representations. Ch. 30
  258. Yuan, X., Chen, Z., Zhang, J., Xiong, H., Ye, N., Li, Y., and Gu, Q. (2026). Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery. ICLR 2026. Ch. 30
  259. Yue, Y., and Joachims, T. (2009). Interactively optimizing information retrieval systems as a dueling bandits problem. Proceedings of the 26th Annual International Conference on Machine Learning. Ch. 26 Ch. 29
  260. Yue, Y., Broder, J., Kleinberg, R., and Joachims, T. (2012). The K-armed Dueling Bandits Problem. Journal of Computer and System Sciences. Ch. 29
  261. Zhang, X. (2025). PABBO code repository: evaluation config evaluate.yaml. GitHub. software Ch. 27 Ch. 30 Ch. 31
  262. Zhang, X. (2026). PABBO. GitHub. software Ch. 31
  263. Zhang, X., Huang, D., Kaski, S., and Martinelli, J. (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Ch. 26 Ch. 27 Ch. 28 Ch. 30 Ch. 31
  264. Zhang, X., Hassan, C., Martinelli, J., Huang, D., and Kaski, S. (2026a). In-Context Multi-Objective Optimization. International Conference on Learning Representations. Ch. 30
  265. Zhang, R., Zhu, X., Pourebadi Khotbehsara, M., Dao, W., Bıyık, E., and Culbertson, H. (2026b). Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback. UMAP 2026 (per Semantic Scholar). Ch. 27
  266. Zhu, M. (2024). Global and preference-based optimization using surrogate-based methods. IMT School for Advanced Studies Lucca. doi:10.13118/imtlucca/e-theses/415. thesis Ch. 31
  267. Zhu, M. (2025). PWAS. GitHub. software Ch. 31
  268. Zhu, M., and Bemporad, A. (2025). Global and Preference-Based Optimization with Mixed Variables Using Piecewise Affine Surrogates. Journal of Optimization Theory and Applications. Ch. 28
  269. Zhu, M., Piga, D., and Bemporad, A. (2022). C-GLISp: Preference-Based Global Optimization Under Unknown Constraints With Applications to Controller Calibration. IEEE Transactions on Control Systems Technology. Ch. 27 Ch. 28 Ch. 31
  270. Zhu, B., Jordan, M., and Jiao, J. (2023). Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. International Conference on Machine Learning. Ch. 29
  271. Ziomek, J., Adachi, M., and Osborne, M. A. (2024). Bayesian Optimisation with Unknown Hyperparameters: Regret Bounds Logarithmically Closer to Optimal. NeurIPS 2024. Ch. 30