Bayesian Optimization
Part VIII: Neighbors in Computing
中文

Neighbors in Computing

Preferential Bayesian optimization shares its likelihood with the reward models used to align large language models, and its questions with reward learning, recommender systems, learning to rank, decision analysis, and automated science. This part maps the traffic between these fields: where one solved a problem first, what moves across, and where an analogy that looks exact breaks down.

The neighbors matter to this book because several of them reached the measurement problem first (Section 45.1). Alignment research measured how far a language-model judge departs from a person, learning to rank learned to correct for the position in which an option is shown, and AI safety formalized systems that change the people they learn from. The first chapter treats large language models in both directions, as components inside a preference loop and as the target of preference learning at a scale no design study reaches. The second visits the other neighbors and collects the problems they solved first that preferential optimization can reuse.

The part assumes Chapter 16 and Chapter 18.

Chapters in this part

  1. 35 Preferences and Large Language Models

    How language-model alignment learns from comparisons, what language models do inside preference loops, which PBO and dueling-bandit methods have been applied to language-model problems and with what results, the Bradley-Terry likelihood both fields share and the seven analogies between them that break, and what has flowed back.

  2. 36 Reward Learning, Recommendation, Ranking, and Automated Science

    Seven computing fields that learn from human judgments, and what they solved first: reward learning from trajectory comparisons, the AI safety analysis of systems that influence the people they measure, decision-theoretic preference elicitation, recommender feedback loops, interactive evolution, learning to rank, and self-driving labs. With an interactive look at position bias and a table of results preferential optimization can adopt.

References for Part VIII

224 works cited across this part's chapters.

  1. Adesiji, A. D., Wang, J., Kuo, C.-S., and Brown, K. A. (2026). Benchmarking self-driving labs. Digital Discovery. Ch. 36
  2. Agnihotri, A., Jain, R., Ramachandran, D., and Wen, Z. (2024). Online Bandit Learning with Offline Preference Data for Improved RLHF. arXiv (not accepted at TMLR). preprint Ch. 35
  3. Ahmed, M. H., and Ghasemi, M. (2026). Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare. Transactions on Machine Learning Research. Ch. 35
  4. Alanazi, E., Mouhoub, M., and Zilles, S. (2020). The complexity of exact learning of acyclic conditional preference networks from swap examples. Artificial Intelligence. doi:10.1016/j.artint.2019.103182. Ch. 36
  5. Amirian, B., Dale, A. S., Kalinin, S., and Hattrick-Simpers, J. (2025). Building Trustworthy AI for Materials Discovery: From Autonomous Laboratories to Z-scores. arXiv. preprint Ch. 36
  6. An, Z., Nakshbandi, D., and Du, W. (2026). Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences. arXiv. preprint Ch. 35
  7. Ananthakrishnan, N., Bedaywi, M., Jordan, M. I., Russell, S., and Haghtalab, N. (2026). Provably Optimal Learning Algorithms for Assistance Games. arXiv (a 2026 AI4GOOD Workshop version also exists). preprint Ch. 36
  8. Antognini, D., and Faltings, B. (2021). Fast Multi-Step Critiquing for VAE-based Recommender Systems. Fifteenth ACM Conference on Recommender Systems. doi:10.1145/3460231.3474249. Ch. 36
  9. Asghari, S. M., Chute, C., Dwaracherla, V., Lu, X., Jafarnia, M., Minden, V., Wen, Z., and Van Roy, B. (2026). Efficient Exploration at Scale. arXiv. preprint Ch. 35
  10. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. Ch. 36
  11. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). Ch. 35
  12. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024b). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. arXiv. preprint Ch. 35
  13. Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R. (2024). A General Theoretical Paradigm to Understand Learning from Human Preferences. AISTATS 2024. Ch. 35
  14. Bai, C., Zhang, Y., Qiu, S., Zhang, Q., Xu, K., and Li, X. (2025). Online Preference Alignment for Language Models via Count-based Exploration. ICLR 2025. Ch. 35
  15. Bakshy, E., Messing, S., and Adamic, L. A. (2015). Exposure to ideologically diverse news and opinion on Facebook. Science. Ch. 36
  16. Blum, A., Jackson, J., Sandholm, T., and Zinkevich, M. (2004). Preference Elicitation and Query Learning. Journal of Machine Learning Research. Ch. 36
  17. Bonilla, E. V., Zhao, H., and Steinberg, D. M. (2026). Causal Preference Elicitation. ICML 2026 (per OpenReview). Ch. 36
  18. Bontrager, P., Lin, W., Togelius, J., and Risi, S. (2018). Deep Interactive Evolution. EvoMUSART 2018. Ch. 36
  19. Bose, A., Xiong, Z., Chi, Y., Du, S. S., Xiao, L., and Fazel, M. (2025). LoRe: Personalizing LLMs via Low-Rank Reward Modeling. Conference on Language Modeling (COLM 2025). Ch. 35
  20. Boutilier, C. (2002). A POMDP Formulation of Preference Elicitation Problems. Proceedings of the Eighteenth National Conference on Artificial Intelligence (AAAI-02). Ch. 36
  21. Bukharin, A., Li, Y., He, P., and Zhao, T. (2023). Deep Reinforcement Learning from Hierarchical Preference Design. International Conference on Machine Learning (ICML 2025). Ch. 36
  22. Cai, H., Li, Y., Yu, T., Zhu, F., Wang, W., Feng, F., and Li, W. (2026). One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment. SIGIR 2026. Ch. 35
  23. Carroll, M., Dragan, A., Russell, S., and Hadfield-Menell, D. (2022). Estimating and Penalizing Induced Preference Shifts in Recommender Systems. ICML 2022. Ch. 36
  24. Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. Ch. 36
  25. Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., … Hadfield-Menell, D. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. TMLR 2023. Ch. 36
  26. Cen, S., Mei, J., Goshvadi, K., Dai, H., Yang, T., Yang, S., … Dai, B. (2025). Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF. ICLR 2025. Ch. 35
  27. Cercola, M., Capretti, V., and Formentin, S. (2026a). Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference. IFAC Journal of Systems and Control. doi:10.1016/j.ifacsc.2026.100398. Ch. 35
  28. Chajewska, U., Koller, D., and Parr, R. (2000). Making Rational Decisions using Adaptive Utility Elicitation. Proceedings of the Seventeenth National Conference on Artificial Intelligence (AAAI-00). Ch. 36
  29. Chaney, A. J. B., Stewart, B. M., and Engelhardt, B. E. (2018). How algorithmic confounding in recommendation systems increases homogeneity and decreases utility. Proceedings of the 12th ACM Conference on Recommender Systems. Ch. 36
  30. Chang, M.-C., Amsler, M., Sutherland, D. R., Ament, S., Gann, K. R., Zhou, L., … Thompson, M. O. (2026). Autonomous Materials Exploration by Integrating Automated Phase Identification and AI-Assisted Human Reasoning. arXiv. preprint Ch. 36
  31. Chawla, A., Thompson, W. H. W., and Young, J.-G. (2026). Multiple latent orderings better predict language model preferences. arXiv. preprint Ch. 35
  32. Chen, A., Malladi, S., Zhang, L. H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K. (2024). Preference Learning Algorithms Do Not Learn Preference Rankings. Advances in Neural Information Processing Systems. doi:10.52202/079017-3234. Ch. 35
  33. Chen, M., Chen, Y., Sun, W., and Zhang, X. (2025a). Avoiding scaling in RLHF through Preference-based Exploration. Advances in Neural Information Processing Systems 38 (NeurIPS 2025). doi:10.52202/085713-5485. Ch. 35
  34. Chen, D., Chen, Y., Rege, A., and Vinayak, R. K. (2025c). PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences. ICLR 2025. Ch. 35
  35. Cheng, J., Xiong, G., Dai, X., Miao, Q., Lv, Y., and Wang, F.-Y. (2024). RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences. ICML 2024. Ch. 36
  36. Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., … Stoica, I. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. ICML 2024. Ch. 35
  37. Chiang, C.-K., Ishida, T., and Sugiyama, M. (2025). LLM Routing with Dueling Feedback. arXiv (not accepted at ICLR 2026). preprint Ch. 35
  38. Chidambaram, K., Seetharaman, K. V., and Syrgkanis, V. (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. Ch. 35
  39. Choi, H., Jung, S., Ahn, H., and Moon, T. (2024). Listwise Reward Estimation for Offline Preference-based Reinforcement Learning. ICML 2024. Ch. 36
  40. Christakopoulou, K., Radlinski, F., and Hofmann, K. (2016). Towards Conversational Recommender Systems. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Ch. 36
  41. Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems. Ch. 35 Ch. 36
  42. Coste, T., Anwar, U., Kirk, R., and Krueger, D. (2024). Reward Model Ensembles Help Mitigate Overoptimization. International Conference on Learning Representations. Ch. 35
  43. Das, N., Chakraborty, S., Pacchiano, A., and Chowdhury, S. R. (2025). Active Preference Optimization for Sample Efficient RLHF. Machine Learning and Knowledge Discovery in Databases. Research Track. doi:10.1007/978-3-032-06096-9_6. Ch. 35
  44. Dean, S., and Morgenstern, J. (2022). Preference Dynamics Under Personalized Recommendations. EC 2022. Ch. 36
  45. Defresne, M., Mandi, J., and Guns, T. (2025). Preference Elicitation for Multi-objective Combinatorial Optimization with Active Learning and Maximum Likelihood Estimation. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. Ch. 36
  46. Deneault, J. R., Kim, W., Kim, J., Gu, Y., Chang, J., Maruyama, B., Myung, J. I., and Pitt, M. A. (2025). Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities. Digital Discovery. Ch. 36
  47. Ding, L., Zhang, J., Clune, J., Spector, L., and Lehman, J. (2024). Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization. ICML 2024. Ch. 36
  48. Du, Z., Zhang, H., Zhu, H., and Zhang, B. (2026). Optimal Design for Active Preference Learning with Biased LLM Judges. arXiv. preprint Ch. 35
  49. Duan, Z., Rong, G., Li, Z., Chen, B., Zhou, M., and Guo, D. (2026). Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling. ICML 2026. Ch. 35
  50. Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., … Hashimoto, T. B. (2023). AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback. Advances in Neural Information Processing Systems. doi:10.52202/075280-1308. Ch. 35
  51. Dwaracherla, V., Asghari, S. M., Hao, B., and Van Roy, B. (2024). Efficient Exploration for LLMs. ICML 2024. Ch. 35
  52. Eichelbeck, M., Voigt, T., and Althoff, M. (2026). Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent Space. International Conference on Learning Representations. Ch. 35
  53. Emmons, S., Oesterheld, C., Conitzer, V., and Russell, S. (2025). Observation Interference in Partially Observable Assistance Games. ICML 2025. Ch. 36
  54. Evans, C., and Kasirzadeh, A. (2023). User Tampering in Reinforcement Learning Recommender Systems. AIES 2023. Ch. 36
  55. Feng, S., and Fu, J. (2025). Thompson Sampling in Online RLHF with General Function Approximation. arXiv. preprint Ch. 35
  56. Fickinger, A., Zhuang, S., Hadfield-Menell, D., and Russell, S. (2020). Multi-Principal Assistance Games. arXiv. preprint Ch. 36
  57. Gao, C., Lei, W., He, X., de Rijke, M., and Chua, T.-S. (2021). Advances and Challenges in Conversational Recommender Systems: A Survey. AI Open. Ch. 36
  58. Gao, L., Schulman, J., and Hilton, J. (2023). Scaling Laws for Reward Model Overoptimization. Proceedings of the 40th International Conference on Machine Learning (ICML 2023). Ch. 36
  59. Gauthier, G., Hodler, R., Widmer, P., and Zhuravskaya, E. (2026). The political effects of X’s feed algorithm. Nature. doi:10.1038/s41586-026-10098-2. Ch. 36
  60. Gharat, S., Karamchandani, N., and Nair, J. (2026). Cost-Aware Best-LLM Identification using Dueling Feedback. Advances in Neural Information Processing Systems. Ch. 35
  61. Gölz, P., Haghtalab, N., and Yang, K. (2025). Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? NeurIPS 2025. Ch. 35
  62. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 35
  63. Gupta, R., Hartford, J., and Liu, B. (2025). LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Findings of the Association for Computational Linguistics: EMNLP 2025. Ch. 35
  64. Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S., and Dragan, A. (2017). Inverse Reward Design. NeurIPS 2017. Ch. 36
  65. Hahami, E., Zimmermann, Y., Zhou, R., and Benarroch Jedlicki, J. (2026). A Unifying Lens on Reward Uncertainty in RLHF. arXiv. preprint Ch. 35
  66. Handa, K., Gal, Y., Pavlick, E., Goodman, N., Andreas, J., Tamkin, A., and Li, B. Z. (2024). Bayesian Preference Elicitation with Language Models. arXiv. preprint Ch. 35
  67. Hatgis-Kessell, S., Knox, W. B., Booth, S., and Stone, P. (2025). Influencing Humans to Conform to Preference Models for RLHF. Transactions on Machine Learning Research. Ch. 36
  68. Heckel, R., Shah, N. B., Ramchandran, K., and Wainwright, M. J. (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. Ch. 36
  69. Hejna, I. D. J., and Sadigh, D. (2022). Few-Shot Preference Learning for Human-in-the-Loop RL. Conference on Robot Learning. Ch. 36
  70. Hejna, J., and Sadigh, D. (2023). Inverse Preference Learning: Preference-based RL without a Reward Function. NeurIPS 2023. Ch. 36
  71. Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D. (2024). Contrastive Preference Learning: Learning from Human Feedback without RL. ICLR 2024. Ch. 36
  72. Herin, M., Perny, P., and Sokolovska, N. (2024). Noise-Tolerant Active Preference Learning for Multicriteria Choice Problems. Algorithmic Decision Theory. Ch. 36
  73. Hong, J., Bhatia, K., and Dragan, A. (2023). On the Sensitivity of Reward Inference to Misspecified Human Models. ICLR 2023. Ch. 36
  74. Hopkins, M., Kane, D., Lovett, S., and Mahajan, G. (2020). Noise-tolerant, Reliable Active Classification with Comparison Queries. Conference on Learning Theory. Ch. 36
  75. Hosseinmardi, H., Ghasemian, A., Rivera-Lanas, M., Horta Ribeiro, M., West, R., and Watts, D. J. (2024). Causally estimating the effect of YouTube’s recommender system using counterfactual bots. Proceedings of the National Academy of Sciences. Ch. 36
  76. Hu, X., Li, J., Zhan, X., Jia, Q.-S., and Zhang, Y.-Q. (2024). Query-Policy Misalignment in Preference-Based Reinforcement Learning. ICLR 2024. Ch. 36
  77. Jannach, D., Manzoor, A., Cai, W., and Chen, L. (2021). A Survey on Conversational Recommender Systems. ACM Computing Surveys. Ch. 36
  78. Ji, K., He, J., and Gu, Q. (2024). Reinforcement Learning from Human Feedback with Active Queries. TMLR. Ch. 35
  79. Joachims, T., Swaminathan, A., and Schnabel, T. (2017). Unbiased Learning-to-Rank with Biased Feedback. WSDM 2017. Ch. 36
  80. Johnston, C. M., Vossler, P., Blessenohl, S., and Vayanos, P. (2023). Deploying a Robust Active Preference Elicitation Algorithm on MTurk: Experiment Design, Interface, and Evaluation for COVID-19 Patient Prioritization. EAAMO 2023. Ch. 36
  81. Kalimeris, D., Bhagat, S., Kalyanaraman, S., and Weinsberg, U. (2021). Preference Amplification in Recommender Systems. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. Ch. 36
  82. Kalinin, S. V., Liu, Y., Biswas, A., Duscher, G., Pratiush, U., Roccapriore, K., Ziatdinov, M., and Vasudevan, R. (2024). Human-in-the-loop: The future of Machine Learning in Automated Electron Microscopy. Microscopy Today. doi:10.1093/mictod/qaad096. Ch. 36
  83. Kane, D. M., Lovett, S., Moran, S., and Zhang, J. (2017). Active classification with comparison queries. FOCS 2017. Ch. 36
  84. Karlekar, S., Zheng, C., Saebo, M., Beltran-Velez, N., Yu, S., Bowlan, J., Kucer, M., and Blei, D. (2026). Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. ICLR 2026 RSI Workshop. workshop paper Ch. 35
  85. Karwowski, J., Hayman, O., Bai, X., Kiendlhofer, K., Griffin, C., and Skalse, J. (2024). Goodhart's Law in Reinforcement Learning. ICLR 2024. Ch. 36
  86. Katkuri, S., Kawada, M., and Wachs, J. (2026). Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning. arXiv. preprint Ch. 35
  87. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. Ch. 35
  88. Kim, G., and Kim, E. (2026). Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback. ICLR 2026. Ch. 35
  89. Kirk, H. R., Leqi, L., Zeng, F., Davidson, H., Vidgen, B., Summerfield, C., and Hale, S. A. (2026). PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users. arXiv. preprint Ch. 35
  90. Kleine Buening, T., Gan, J., Mandal, D., and Kwiatkowska, M. (2025). Strategyproof Reinforcement Learning from Human Feedback. NeurIPS 2025. Ch. 36
  91. Knox, W. B., Hatgis-Kessell, S., Booth, S., Niekum, S., Stone, P., and Allievi, A. (2024). Models of human preference for learning reward functions. TMLR 2024. Ch. 36
  92. Kobalczyk, K., Astorga, N., Liu, T., and van der Schaar, M. (2025). Active Task Disambiguation with LLMs. ICLR 2025. Ch. 35
  93. Kobalczyk, K., Lin, Z. J., Letham, B., Zhao, Z., Balandat, M., and Bakshy, E. (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. Ch. 35
  94. Korbak, T., Perez, E., and Buckley, C. L. (2022). RL with KL penalties is better viewed as Bayesian inference. Findings of the Association for Computational Linguistics: EMNLP 2022. doi:10.18653/v1/2022.findings-emnlp.77. Ch. 35
  95. Kristiadi, A., Strieth-Kalthoff, F., Skreta, M., Poupart, P., Aspuru-Guzik, A., and Pleiss, G. (2024a). A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules? International Conference on Machine Learning. Ch. 35
  96. Kuric, E., Demcak, P., and Krajcovic, M. (2026). Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design? arXiv. preprint Ch. 35
  97. Kutulakos, Z., and Slade, P. (2024). Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance. bioRxiv. preprint Ch. 36
  98. Kveton, B., Li, X., McAuley, J., Rossi, R., Shang, J., Wu, J., and Yu, T. (2025). Active Learning for Direct Preference Optimization. arXiv. preprint Ch. 35
  99. Kwa, T., Thomas, D., and Garriga-Alonso, A. (2024). Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification. NeurIPS 2024. Ch. 36
  100. Laidlaw, C., Bronstein, E., Guo, T., Feng, D., Berglund, L., Svegliato, J., Russell, S., and Dragan, A. (2025). AssistanceZero: Scalably Solving Assistance Games. ICML 2025. Ch. 36
  101. Lang, L., Foote, D., Russell, S., Dragan, A., Jenner, E., and Emmons, S. (2024). When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback. NeurIPS 2024. Ch. 36
  102. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. Ch. 35
  103. Lee, K., Smith, L., Dragan, A., and Abbeel, P. (2021a). B-Pref: Benchmarking Preference-Based Reinforcement Learning. NeurIPS 2021 Datasets and Benchmarks. Ch. 36
  104. Lee, K., Smith, L., and Abbeel, P. (2021b). PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training. ICML 2021. Ch. 36
  105. Lee, U. H., Shetty, V. S., Franks, P. W., Tan, J., Evangelopoulos, G., Ha, S., and Rouse, E. J. (2023). User preference optimization for control of ankle exoskeletons using sample efficient active learning. Science Robotics. Ch. 36
  106. Li, X., Zhao, H., and Gu, Q. (2024b). Feel-Good Thompson Sampling for Contextual Dueling Bandits. International Conference on Machine Learning. Ch. 35
  107. Li, B. Z., Tamkin, A., Goodman, N., and Andreas, J. (2025b). Eliciting Human Preferences with Language Models. ICLR 2025. Ch. 35
  108. Li, W., Oh, C., and Li, S. (2026d). General Exploratory Bonus for Optimistic Exploration in RLHF. International Conference on Learning Representations. Ch. 35
  109. Li, Z., Liao, Y.-C., and Holz, C. (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. Ch. 35
  110. Liao, Y.-C., Belo, J., Moon, H.-S., Steimle, J., and Feit, A. M. (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. Ch. 35
  111. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. Ch. 35
  112. Lin, Y., Seto, S., ter Hoeve, M., Metcalf, K., Theobald, B.-J., Wang, X., … Zhang, T. (2024a). On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization. Findings of the Association for Computational Linguistics: EMNLP 2024. doi:10.18653/v1/2024.findings-emnlp.940. Ch. 35
  113. Lin, X., Dai, Z., Verma, A., Ng, S.-K., Jaillet, P., and Low, B. K. H. (2024b). Prompt Optimization with Human Feedback. ICML 2024 MHFAIA Workshop (no formal proceedings). workshop paper Ch. 35
  114. Lindner, D., Turchetta, M., Tschiatschek, S., Ciosek, K., and Krause, A. (2021). Information Directed Reward Learning for Reinforcement Learning. NeurIPS 2021. Ch. 36
  115. Liu, T., Astorga, N., Seedat, N., and van der Schaar, M. (2024a). Large Language Models to Enhance Bayesian Optimization. International Conference on Learning Representations. Ch. 35
  116. Liu, Z., Chen, C., Du, C., Lee, W. S., and Lin, M. (2024b). Sample-Efficient Alignment for LLMs. NeurIPS 2024 LanGame Workshop. workshop paper Ch. 35
  117. Liu, T., Qin, Z., Wu, J., Shen, J., Khalman, M., Joshi, R., … Wang, X. (2025a). LiPO: Listwise Preference Optimization through Learning-to-Rank. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). doi:10.18653/v1/2025.naacl-long.121. Ch. 35
  118. Liu, N., Hu, X. E., Savas, Y., Baum, M. A., Berinsky, A. J., Chaney, A. J. B., … Stewart, B. M. (2025b). Short-term exposure to filter-bubble recommendation systems has limited polarization effects: Naturalistic experiments on YouTube. Proceedings of the National Academy of Sciences. Ch. 36
  119. Liu, N., Sun, C., Klinkner, K., and Malmasi, S. (2026a). Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph. arXiv. preprint Ch. 35
  120. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. Ch. 35
  121. Lou, X., Yan, D., Shen, W., Yan, Y., Xie, J., and Zhang, J. (2024). Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown. arXiv (withdrawn from ICLR 2025). preprint Ch. 35
  122. Ma, Q., Gao, D., Cai, R., Zhao, B., Zhou, H., Zhang, J., and Zhao, Z. (2026). Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization. COLM 2026. Ch. 35
  123. Mahmud, S., Nakamura, M., and Zilberstein, S. (2025). MAPLE: A Framework for Active Preference Learning Guided by Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v39i26.34964. Ch. 35
  124. Mansoury, M., Abdollahpouri, H., Pechenizkiy, M., Mobasher, B., and Burke, R. (2020). Feedback Loop and Bias Amplification in Recommender Systems. CIKM 2020. Ch. 36
  125. Martin, R. M., and Collins, S. H. (2026). Improving CMA-ES Convergence Speed, Efficiency, and Reliability in Noisy Robot Optimization Problems. Evolutionary Computation. Ch. 36
  126. Martin, C., Boutilier, C., Meshi, O., and Sandholm, T. (2024). Model-Free Preference Elicitation. Thirty-Third International Joint Conference on Artificial Intelligence. Ch. 36
  127. McElfresh, D. C., Chan, L., Doyle, K., Sinnott-Armstrong, W., Conitzer, V., Schaich Borg, J., and Dickerson, J. P. (2021). Indecision Modeling. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v35i7.16746. Ch. 36
  128. Mehta, V., Belakaria, S., Das, V., Neopane, O., Dai, Y., Bogunovic, I., … Neiswanger, W. (2025). Sample Efficient Preference Alignment in LLMs via Active Exploration. COLM 2025. Ch. 35
  129. Melikidze, D., Schneider, M., Lam, J., Wertich, M., Hakimi, I., Pásztor, B., and Krause, A. (2026). ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning. ICML 2026. Ch. 35
  130. Melo, L. C., Tigas, P., Abate, A., and Gal, Y. (2024). Deep Bayesian Active Learning for Preference Modeling in Large Language Models. NeurIPS 2024. Ch. 35
  131. Menn, J., Kober, M., Brunzema, P., Stenger, D., and Trimpe, S. (2026a). Local Preferential Bayesian Optimization. arXiv. preprint Ch. 35
  132. Meta Platforms, Inc. (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. software Ch. 35
  133. Meta Platforms, Inc. (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. software Ch. 35
  134. Mikkola, P., Todorović, M., Järvi, J., Rinke, P., and Kaski, S. (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 36
  135. Muldrew, W., Hayes, P., Zhang, M., and Barber, D. (2024). Active Preference Learning for Large Language Models. ICML 2024. Ch. 35
  136. Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., … Piot, B. (2024). Nash Learning from Human Feedback. ICML 2024. Ch. 35
  137. Nan, T., Li, X., Kroer, C., and Lin, T. (2026). Efficient Exploration for Iterative Nash Preference Optimization. arXiv. preprint Ch. 35
  138. Nguyen, S., Liu, X., and Senanayake, R. (2026). CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM. International Conference on Machine Learning (ICML 2026). Ch. 35
  139. Nisan, N., and Segal, I. (2006). The communication requirements of efficient allocations and supporting prices. Journal of Economic Theory. Ch. 36
  140. Niwa, R., Yoshida, S., Koyama, Y., and Ushiku, Y. (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. Ch. 35
  141. Noothigattu, R., Peters, D., and Procaccia, A. (2020). Axioms for Learning from Pairwise Comparisons. Advances in Neural Information Processing Systems. Ch. 36
  142. Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J. (2020). Dueling Posterior Sampling for Preference-Based Reinforcement Learning. Conference on Uncertainty in Artificial Intelligence. Ch. 36
  143. Oh, G., Lee, J., Park, J., Yu, Y., Bae, W., and Noh, J. (2026b). Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs. ICLR 2026 Workshop: I Can't Believe It's Not Better (ICBINB). workshop paper Ch. 35
  144. Oko, K., Ulichney, A., Haghtalab, N., and Bao, H. (2026). Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner. International Conference on Machine Learning. Ch. 35
  145. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems. Ch. 35
  146. Pan, A., Bhatia, K., and Steinhardt, J. (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. ICLR 2022. Ch. 36
  147. Panickssery, A., Bowman, S. R., and Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS 2024. Ch. 35
  148. Papadimitriou, C. H., and Tsitsiklis, J. N. (1987). The Complexity of Markov Decision Processes. Mathematics of Operations Research. Ch. 36
  149. Pásztor, B., Kassraie, P., and Krause, A. (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. Ch. 35
  150. Peng, Y.-H., Bigham, J. P., and Wu, J. (2026). Efficient Personalization of Generative User Interfaces. arXiv. preprint Ch. 35
  151. Piriyakulkij, W. T., Kuleshov, V., and Ellis, K. (2023). Active Preference Inference using Language Models and Probabilistic Reasoning. NeurIPS 2023 FMDM Workshop. workshop paper Ch. 35
  152. Poddar, S., Wan, Y., Ivison, H., Gupta, A., and Jaques, N. (2024). Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning. Advances in Neural Information Processing Systems 37 (NeurIPS 2024). doi:10.52202/079017-1664. Ch. 35
  153. Pratiush, U., Roccapriore, K. M., Liu, Y., Duscher, G., Ziatdinov, M., and Kalinin, S. V. (2025). Building Workflows for Interactive Human in the Loop Automated Experiment (hAE) in STEM-EELS. Digital Discovery. doi:10.1039/d5dd00033e. Ch. 36
  154. Pukdee, R., Balcan, M.-F., and Ravikumar, P. (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. Ch. 36
  155. Qiu, L., Sha, F., Allen, K., Kim, Y., Linzen, T., and van Steenkiste, S. (2026). Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models. Nature Communications. doi:10.1038/s41467-025-67998-6. Ch. 35
  156. Qu, Z., Zhang, M., Kong, M., Li, X., Shang, Z., Wang, Z., … Dai, Z. (2026). T-POP: Test-Time Personalization with Online Preference Feedback. International Conference on Machine Learning. Ch. 35
  157. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. Ch. 35
  158. Rafailov, R., Chittepu, Y., Park, R., Sikchi, H., Hejna, J., Knox, B., Finn, C., and Niekum, S. (2024). Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms. NeurIPS 2024. Ch. 36
  159. Rajagopalan, R., Dutta, D., Wei, Y.-L., and Roy Choudhury, R. (2026). Personalized Image Generation via Human-in-the-loop Bayesian Optimization. International Conference on Machine Learning. Ch. 35
  160. Ramos, M. C., Michtavy, S. S., Porosoff, M. D., and White, A. D. (2026). Bayesian Optimization of Catalysis With In-Context Learning. ACS Central Science. doi:10.1021/acscentsci.5c02418. Ch. 35
  161. Ranković, B., Griffiths, R.-R., and Schwaller, P. (2026). Large language models as uncertainty-calibrated optimizers for experimental discovery. Nature Machine Intelligence. doi:10.1038/s42256-026-01283-z. Ch. 35
  162. Rodrigues, C., Vas, O., DCosta, I. A., and Prabhakaran, N. K. (2026). When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model. arXiv. preprint Ch. 35
  163. Rychert, A., Spagnolo, G., and Posashkov, E. (2025). Reproducibility Study of Large Language Model Bayesian Optimization. arXiv. preprint Ch. 35
  164. Saha, A., Pacchiano, A., and Lee, J. (2023). Dueling RL: Reinforcement Learning with Trajectory Preferences. International Conference on Artificial Intelligence and Statistics. Ch. 36
  165. Saracay, I., Schmidt, L., and Guestrin, C. (2026). Beyond expert users: agents should help users construct preferences, not just elicit them. Conference on Language Modeling (COLM 2026). Ch. 35
  166. Scheid, A., Boursier, E., Durmus, A., Jordan, M. I., Ménard, P., Moulines, E., and Valko, M. (2024). Optimal Design for Reward Modeling in RLHF. arXiv. preprint Ch. 35
  167. Schoinas, E., Rastogi, A., Carter, A., Granley, J., and Beyeler, M. (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. Ch. 35
  168. Seshadri, P., Cahyawijaya, S., Odumakinde, A., Singh, S., and Goldfarb-Tarrant, S. (2026). Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2026.acl-long.2192. Ch. 35
  169. Sfikas, K., Liapis, A., and Yannakakis, G. N. (2023). Controllable Exploration of a Design Space via Interactive Quality Diversity. arXiv (parts published at GECCO 2023). preprint Ch. 36
  170. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. J. (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. Ch. 36
  171. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. preprint Ch. 36
  172. Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., … Perez, E. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. Ch. 36
  173. Shen, Y., Sun, H., and Ton, J.-F. (2025a). Active Reward Modeling: Adaptive Preference Labeling for Large Language Model Alignment. International Conference on Machine Learning. Ch. 35
  174. Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., and Vosoughi, S. (2025). Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics. Ch. 35
  175. Shields, B. J., Stevens, J., Li, J., Parasram, M., Damani, F., Alvarado, J. I. M., … Doyle, A. G. (2021). Bayesian reaction optimization as a tool for chemical synthesis. Nature. Ch. 36
  176. Singh, U., Chakraborty, S., Suttle, W. A., Sadler, B. M., Asher, D. E., Sahu, A. K., … Bedi, A. S. (2024). Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach. International Conference on Learning Representations (ICLR 2026). Ch. 36
  177. Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. Ch. 35
  178. Skalse, J., Howe, N. H. R., Krasheninnikov, D., and Krueger, D. (2022). Defining and Characterizing Reward Hacking. Advances in Neural Information Processing Systems 35 (NeurIPS 2022). Ch. 36
  179. Skalse, J., Farrugia-Roberts, M., Russell, S., Abate, A., and Gleave, A. (2023). Invariance in Policy Optimisation and Partial Identifiability in Reward Learning. ICML 2023. Ch. 36
  180. Smith, J. E., and Winkler, R. L. (2006). The Optimizer’s Curse: Skepticism and Postdecision Surprise in Decision Analysis. Management Science. Ch. 36
  181. Song, Y., Swamy, G., Singh, A., Bagnell, J. A., and Sun, W. (2024). The Importance of Online Data: Understanding Preference Fine-tuning via Coverage. NeurIPS 2024. Ch. 35
  182. Sun, H., Shen, Y., and Ton, J.-F. (2025). Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives. ICLR 2025. Ch. 35
  183. Surana, R., Li, X., Yu, S., Shen, Y. J., Wang, C., Yu, T., … Wu, J. (2026). MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization. arXiv. preprint Ch. 35
  184. Takagi, H. (2001). Interactive evolutionary computation: fusion of the capabilities of EC optimization and human evaluation. Proceedings of the IEEE. Ch. 36
  185. Takagi, H., and Pallez, D. (2009). Paired Comparisons-based Interactive Differential Evolution. NaBIC 2009. Ch. 36
  186. Takeno, S., Nomura, M., and Karasuyama, M. (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. Ch. 36
  187. Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., … Piot, B. (2024). Generalized Preference Optimization: A Unified Approach to Offline Alignment. International Conference on Machine Learning. Ch. 35
  188. Thies, S. M. A. R., Bengs, V., Kaufmann, T., Vollmer, S. J., and Hüllermeier, E. (2026a). Calibrated Preference Learning: The Case of Label Ranking. International Conference on Machine Learning (ICML 2026). Ch. 35 Ch. 36
  189. Thies, S. M. A. R., Alfaro, J. C., and Bengs, V. (2026b). MORE-PLR: multi-output regression employed for partial label ranking. Machine Learning 115. Ch. 36
  190. Vayanos, P., Ye, Y., McElfresh, D., Dickerson, J., and Rice, E. (2020). Robust Active Preference Elicitation. arXiv (journal version not found). preprint Ch. 36
  191. Vendrov, I., Lu, T., Huang, Q., and Boutilier, C. (2020). Gradient-based Optimization for Bayesian Preference Elicitation. AAAI 2020. Ch. 36
  192. Verma, A., Dai, Z., Lin, X., Jaillet, P., and Low, B. K. H. (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. Ch. 35
  193. Viappiani, P., and Boutilier, C. (2010). Optimal Bayesian Recommendation Sets and Myopically Optimal Choice Query Sets. Advances in Neural Information Processing Systems. Ch. 36
  194. Viappiani, P., and Boutilier, C. (2020). On the equivalence of optimal recommendation sets and myopically optimal query sets. Artificial Intelligence. Ch. 36
  195. Wang, T., and Boutilier, C. (2003). Incremental Utility Elicitation with the Minimax Regret Decision Criterion. Proceedings of the Eighteenth International Joint Conference on Artificial Intelligence (IJCAI-03). Ch. 36
  196. Wang, Y., and Pei, Y. (2024). A comprehensive survey on interactive evolutionary computation in the first two decades of the 21st century. Applied Soft Computing. doi:10.1016/j.asoc.2024.111950. Ch. 36
  197. Wang, Y., Liu, Q., and Jin, C. (2023a). Is RLHF More Difficult than Standard RL? NeurIPS 2023. Ch. 35 Ch. 36
  198. Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., … Sui, Z. (2024a). Large Language Models are not Fair Evaluators. ACL 2024. Ch. 35 Ch. 36
  199. Wang, Y., Sun, Z., Zhang, J., Xian, Z., Biyik, E., Held, D., and Erickson, Z. (2024b). RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback. International Conference on Machine Learning. Ch. 35
  200. Wang, H., Branke, J., and Poloczek, M. (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. Ch. 36
  201. Wen, J., Zhong, R., Khan, A., Perez, E., Steinhardt, J., Huang, M., … Feng, S. (2025). Language Models Learn to Mislead Humans via RLHF. ICLR 2025. Ch. 36
  202. Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., and Dragan, A. (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. Ch. 36
  203. Won, Y., Lee, H., Hwang, H., and Seo, M. (2025). Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization. arXiv. preprint Ch. 35
  204. Wu, R., and Sun, W. (2024). Making RL with Preference-based Feedback Efficient via Randomization. ICLR 2024. Ch. 35
  205. Wu, Y., Verma, S., Lee, J., Xiong, F., Zhang, P., Awadelkarim, A., … Hill, S. (2026). LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization. Findings of the Association for Computational Linguistics: ACL 2026. doi:10.18653/v1/2026.findings-acl.490. Ch. 35
  206. Xia, F., Liu, H., Yue, Y., and Li, T. (2025). Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of the Association for Computational Linguistics: ACL 2025. doi:10.18653/v1/2025.findings-acl.519. Ch. 35
  207. Xie, T., Foster, D. J., Krishnamurthy, A., Rosset, C., Awadallah, A., and Rakhlin, A. (2025). Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF. ICLR 2025. Ch. 35
  208. Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T. (2024). Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint. ICML 2024. Ch. 35
  209. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. Ch. 35
  210. Xu, Y., Ruis, L., Rocktäschel, T., and Kirk, R. (2025a). Investigating Non-Transitivity in LLM-as-a-Judge. International Conference on Machine Learning. Ch. 35
  211. Yang, A. X., Robeyns, M., Coste, T., Shi, Z., Wang, J., Bou-Ammar, H., and Aitchison, L. (2024). Bayesian Reward Models for LLM Alignment. ICLR 2024 SeT LLM Workshop; ICML 2024 SPIGM Workshop. workshop paper Ch. 35
  212. Yang, J., Hu, Z., Qiu, C., Deng, Z., Jiao, X., and Zhou, T. (2026a). Quantifying and Mitigating Self-Preference Bias of LLM Judges. arXiv. preprint Ch. 35
  213. Yang, D., Stante, S., Redhardt, F., Libon, L., Kassraie, P., Hakimi, I., Pásztor, B., and Krause, A. (2026b). RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models. EurIPS 2025 EIML Workshop. workshop paper Ch. 35
  214. Yuan, Y., Hao, J., Ma, Y., Dong, Z., Liang, H., Liu, J., … Zheng, Y. (2024). Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback. ICLR 2024. Ch. 36
  215. Yuan, X., Chen, Z., Zhang, J., Xiong, H., Ye, N., Li, Y., and Gu, Q. (2026). Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery. ICLR 2026. Ch. 35
  216. Zhang, Z.-Y., Han, S., Yao, H., Niu, G., and Sugiyama, M. (2024a). Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought. International Conference on Machine Learning. Ch. 35
  217. Zhang, S., Yu, D., Sharma, H., Zhong, H., Liu, Z., Yang, Z., … Wang, Z. (2024b). Self-Exploring Language Models: Active Preference Elicitation for Online Alignment. TMLR. Ch. 35
  218. Zhang, X., Huang, D., Kaski, S., and Martinelli, J. (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. Ch. 35
  219. Zhao, H., Ye, C., Gu, Q., and Zhang, T. (2025). Sharp Analysis for KL-Regularized Contextual Bandits and RLHF. NeurIPS 2025. Ch. 35
  220. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., … Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems. doi:10.52202/075280-2020. Ch. 35
  221. Zhu, B., Jordan, M., and Jiao, J. (2023). Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. International Conference on Machine Learning. Ch. 35
  222. Zhu, L., Huang, X., and Sang, J. (2024). How Reliable is Your Simulator? Analysis on the Limitations of Current LLM-based User Simulators for Conversational Recommendation. Companion Proceedings of the ACM Web Conference 2024. doi:10.1145/3589335.3651955. workshop paper Ch. 35
  223. Zhuang, S., and Hadfield-Menell, D. (2020). Consequences of Misaligned AI. NeurIPS 2020. Ch. 36
  224. Zintgraf, L. M., Roijers, D. M., Linders, S., Jonker, C. M., and Nowé, A. (2018). Ordered Preference Elicitation Strategies for Supporting Multi-Objective Decision Making. AAMAS 2018. Ch. 36