贝叶斯优化
第八部分:相邻计算领域
EN

相邻计算领域

偏好贝叶斯优化与对齐大语言模型所用的奖励模型共用同一似然,与奖励学习、推荐系统、排序学习、决策分析和自动化科学则面对相同的问题。本部分梳理这些领域之间的往来:某个问题由哪个领域率先解决,哪些内容可以跨领域迁移,看似完全对应的类比又在何处失效。

这些相邻领域对本书的意义在于,其中几个领域更早遇到了测量问题(第 45.1 节)。对齐研究测量了语言模型评判者与人的偏离程度;排序学习学会了校正选项展示位置带来的影响;人工智能安全研究则形式化地刻画了这样一类系统:它们从人身上学习,又反过来改变这些人。第一章从两个方向讨论大语言模型:一是作为偏好回路中的组件,二是作为偏好学习的对象,其规模远超任何设计研究。第二章考察其余相邻领域,汇集它们率先解决、其解法可供偏好优化复用的问题。

本部分以第 16 章与第 18 章为前提。

本部分各章

  1. 35 偏好与大语言模型

    语言模型对齐如何从比较中学习;语言模型在偏好回路中承担哪些角色;哪些偏好贝叶斯优化与对决赌博机方法已用于语言模型问题,效果如何;两个领域共用的 Bradley-Terry 似然与二者之间七个失效的类比;以及哪些成果反过来流回偏好贝叶斯优化。

  2. 36 奖励学习、推荐、排序与自动化科学

    七个从人的判断中学习的计算领域及其率先解决的问题:基于轨迹比较的奖励学习、人工智能安全对会影响其测量对象的系统所做的分析、决策论中的偏好引出、推荐系统的反馈回路、交互式进化、排序学习和自驱动实验室。另附对位置偏差的交互式考察,以及一张可供偏好优化采用的结果表。

第八部分参考文献

本部分各章共引用 224 篇文献。

  1. Adesiji, A. D., Wang, J., Kuo, C.-S., and Brown, K. A. (2026). Benchmarking self-driving labs. Digital Discovery. 第 36 章
  2. Agnihotri, A., Jain, R., Ramachandran, D., and Wen, Z. (2024). Online Bandit Learning with Offline Preference Data for Improved RLHF. arXiv (not accepted at TMLR). 预印本 第 35 章
  3. Ahmed, M. H., and Ghasemi, M. (2026). Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare. Transactions on Machine Learning Research. 第 35 章
  4. Alanazi, E., Mouhoub, M., and Zilles, S. (2020). The complexity of exact learning of acyclic conditional preference networks from swap examples. Artificial Intelligence. doi:10.1016/j.artint.2019.103182. 第 36 章
  5. Amirian, B., Dale, A. S., Kalinin, S., and Hattrick-Simpers, J. (2025). Building Trustworthy AI for Materials Discovery: From Autonomous Laboratories to Z-scores. arXiv. 预印本 第 36 章
  6. An, Z., Nakshbandi, D., and Du, W. (2026). Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences. arXiv. 预印本 第 35 章
  7. Ananthakrishnan, N., Bedaywi, M., Jordan, M. I., Russell, S., and Haghtalab, N. (2026). Provably Optimal Learning Algorithms for Assistance Games. arXiv (a 2026 AI4GOOD Workshop version also exists). 预印本 第 36 章
  8. Antognini, D., and Faltings, B. (2021). Fast Multi-Step Critiquing for VAE-based Recommender Systems. Fifteenth ACM Conference on Recommender Systems. doi:10.1145/3460231.3474249. 第 36 章
  9. Asghari, S. M., Chute, C., Dwaracherla, V., Lu, X., Jafarnia, M., Minden, V., Wen, Z., and Van Roy, B. (2026). Efficient Exploration at Scale. arXiv. 预印本 第 35 章
  10. Astudillo, R., Lin, Z. J., Bakshy, E., and Frazier, P. (2023). qEUBO: A Decision-Theoretic Acquisition Function for Preferential Bayesian Optimization. International Conference on Artificial Intelligence and Statistics. 第 36 章
  11. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024a). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. RecSys 2024 (arXiv v2). 第 35 章
  12. Austin, D. E., Korikov, A., Toroghi, A., and Sanner, S. (2024b). Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation. arXiv. 预印本 第 35 章
  13. Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R. (2024). A General Theoretical Paradigm to Understand Learning from Human Preferences. AISTATS 2024. 第 35 章
  14. Bai, C., Zhang, Y., Qiu, S., Zhang, Q., Xu, K., and Li, X. (2025). Online Preference Alignment for Language Models via Count-based Exploration. ICLR 2025. 第 35 章
  15. Bakshy, E., Messing, S., and Adamic, L. A. (2015). Exposure to ideologically diverse news and opinion on Facebook. Science. 第 36 章
  16. Blum, A., Jackson, J., Sandholm, T., and Zinkevich, M. (2004). Preference Elicitation and Query Learning. Journal of Machine Learning Research. 第 36 章
  17. Bonilla, E. V., Zhao, H., and Steinberg, D. M. (2026). Causal Preference Elicitation. ICML 2026 (per OpenReview). 第 36 章
  18. Bontrager, P., Lin, W., Togelius, J., and Risi, S. (2018). Deep Interactive Evolution. EvoMUSART 2018. 第 36 章
  19. Bose, A., Xiong, Z., Chi, Y., Du, S. S., Xiao, L., and Fazel, M. (2025). LoRe: Personalizing LLMs via Low-Rank Reward Modeling. Conference on Language Modeling (COLM 2025). 第 35 章
  20. Boutilier, C. (2002). A POMDP Formulation of Preference Elicitation Problems. Proceedings of the Eighteenth National Conference on Artificial Intelligence (AAAI-02). 第 36 章
  21. Bukharin, A., Li, Y., He, P., and Zhao, T. (2023). Deep Reinforcement Learning from Hierarchical Preference Design. International Conference on Machine Learning (ICML 2025). 第 36 章
  22. Cai, H., Li, Y., Yu, T., Zhu, F., Wang, W., Feng, F., and Li, W. (2026). One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment. SIGIR 2026. 第 35 章
  23. Carroll, M., Dragan, A., Russell, S., and Hadfield-Menell, D. (2022). Estimating and Penalizing Induced Preference Shifts in Recommender Systems. ICML 2022. 第 36 章
  24. Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024). AI Alignment with Changing and Influenceable Reward Functions. International Conference on Machine Learning. 第 36 章
  25. Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., … Hadfield-Menell, D. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. TMLR 2023. 第 36 章
  26. Cen, S., Mei, J., Goshvadi, K., Dai, H., Yang, T., Yang, S., … Dai, B. (2025). Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF. ICLR 2025. 第 35 章
  27. Cercola, M., Capretti, V., and Formentin, S. (2026a). Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference. IFAC Journal of Systems and Control. doi:10.1016/j.ifacsc.2026.100398. 第 35 章
  28. Chajewska, U., Koller, D., and Parr, R. (2000). Making Rational Decisions using Adaptive Utility Elicitation. Proceedings of the Seventeenth National Conference on Artificial Intelligence (AAAI-00). 第 36 章
  29. Chaney, A. J. B., Stewart, B. M., and Engelhardt, B. E. (2018). How algorithmic confounding in recommendation systems increases homogeneity and decreases utility. Proceedings of the 12th ACM Conference on Recommender Systems. 第 36 章
  30. Chang, M.-C., Amsler, M., Sutherland, D. R., Ament, S., Gann, K. R., Zhou, L., … Thompson, M. O. (2026). Autonomous Materials Exploration by Integrating Automated Phase Identification and AI-Assisted Human Reasoning. arXiv. 预印本 第 36 章
  31. Chawla, A., Thompson, W. H. W., and Young, J.-G. (2026). Multiple latent orderings better predict language model preferences. arXiv. 预印本 第 35 章
  32. Chen, A., Malladi, S., Zhang, L. H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K. (2024). Preference Learning Algorithms Do Not Learn Preference Rankings. Advances in Neural Information Processing Systems. doi:10.52202/079017-3234. 第 35 章
  33. Chen, M., Chen, Y., Sun, W., and Zhang, X. (2025a). Avoiding scaling in RLHF through Preference-based Exploration. Advances in Neural Information Processing Systems 38 (NeurIPS 2025). doi:10.52202/085713-5485. 第 35 章
  34. Chen, D., Chen, Y., Rege, A., and Vinayak, R. K. (2025c). PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences. ICLR 2025. 第 35 章
  35. Cheng, J., Xiong, G., Dai, X., Miao, Q., Lv, Y., and Wang, F.-Y. (2024). RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences. ICML 2024. 第 36 章
  36. Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., … Stoica, I. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. ICML 2024. 第 35 章
  37. Chiang, C.-K., Ishida, T., and Sugiyama, M. (2025). LLM Routing with Dueling Feedback. arXiv (not accepted at ICLR 2026). 预印本 第 35 章
  38. Chidambaram, K., Seetharaman, K. V., and Syrgkanis, V. (2026). Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences. International Conference on Artificial Intelligence and Statistics. 第 35 章
  39. Choi, H., Jung, S., Ahn, H., and Moon, T. (2024). Listwise Reward Estimation for Offline Preference-based Reinforcement Learning. ICML 2024. 第 36 章
  40. Christakopoulou, K., Radlinski, F., and Hofmann, K. (2016). Towards Conversational Recommender Systems. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 第 36 章
  41. Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems. 第 35 章 第 36 章
  42. Coste, T., Anwar, U., Kirk, R., and Krueger, D. (2024). Reward Model Ensembles Help Mitigate Overoptimization. International Conference on Learning Representations. 第 35 章
  43. Das, N., Chakraborty, S., Pacchiano, A., and Chowdhury, S. R. (2025). Active Preference Optimization for Sample Efficient RLHF. Machine Learning and Knowledge Discovery in Databases. Research Track. doi:10.1007/978-3-032-06096-9_6. 第 35 章
  44. Dean, S., and Morgenstern, J. (2022). Preference Dynamics Under Personalized Recommendations. EC 2022. 第 36 章
  45. Defresne, M., Mandi, J., and Guns, T. (2025). Preference Elicitation for Multi-objective Combinatorial Optimization with Active Learning and Maximum Likelihood Estimation. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. 第 36 章
  46. Deneault, J. R., Kim, W., Kim, J., Gu, Y., Chang, J., Maruyama, B., Myung, J. I., and Pitt, M. A. (2025). Preferential Bayesian optimization improves the efficiency of printing objects with subjective qualities. Digital Discovery. 第 36 章
  47. Ding, L., Zhang, J., Clune, J., Spector, L., and Lehman, J. (2024). Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization. ICML 2024. 第 36 章
  48. Du, Z., Zhang, H., Zhu, H., and Zhang, B. (2026). Optimal Design for Active Preference Learning with Biased LLM Judges. arXiv. 预印本 第 35 章
  49. Duan, Z., Rong, G., Li, Z., Chen, B., Zhou, M., and Guo, D. (2026). Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling. ICML 2026. 第 35 章
  50. Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., … Hashimoto, T. B. (2023). AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback. Advances in Neural Information Processing Systems. doi:10.52202/075280-1308. 第 35 章
  51. Dwaracherla, V., Asghari, S. M., Hao, B., and Van Roy, B. (2024). Efficient Exploration for LLMs. ICML 2024. 第 35 章
  52. Eichelbeck, M., Voigt, T., and Althoff, M. (2026). Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent Space. International Conference on Learning Representations. 第 35 章
  53. Emmons, S., Oesterheld, C., Conitzer, V., and Russell, S. (2025). Observation Interference in Partially Observable Assistance Games. ICML 2025. 第 36 章
  54. Evans, C., and Kasirzadeh, A. (2023). User Tampering in Reinforcement Learning Recommender Systems. AIES 2023. 第 36 章
  55. Feng, S., and Fu, J. (2025). Thompson Sampling in Online RLHF with General Function Approximation. arXiv. 预印本 第 35 章
  56. Fickinger, A., Zhuang, S., Hadfield-Menell, D., and Russell, S. (2020). Multi-Principal Assistance Games. arXiv. 预印本 第 36 章
  57. Gao, C., Lei, W., He, X., de Rijke, M., and Chua, T.-S. (2021). Advances and Challenges in Conversational Recommender Systems: A Survey. AI Open. 第 36 章
  58. Gao, L., Schulman, J., and Hilton, J. (2023). Scaling Laws for Reward Model Overoptimization. Proceedings of the 40th International Conference on Machine Learning (ICML 2023). 第 36 章
  59. Gauthier, G., Hodler, R., Widmer, P., and Zhuravskaya, E. (2026). The political effects of X’s feed algorithm. Nature. doi:10.1038/s41586-026-10098-2. 第 36 章
  60. Gharat, S., Karamchandani, N., and Nair, J. (2026). Cost-Aware Best-LLM Identification using Dueling Feedback. Advances in Neural Information Processing Systems. 第 35 章
  61. Gölz, P., Haghtalab, N., and Yang, K. (2025). Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? NeurIPS 2025. 第 35 章
  62. González, J., Dai, Z., Damianou, A., and Lawrence, N. D. (2017). Preferential Bayesian Optimization. International Conference on Machine Learning. 第 35 章
  63. Gupta, R., Hartford, J., and Liu, B. (2025). LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Findings of the Association for Computational Linguistics: EMNLP 2025. 第 35 章
  64. Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S., and Dragan, A. (2017). Inverse Reward Design. NeurIPS 2017. 第 36 章
  65. Hahami, E., Zimmermann, Y., Zhou, R., and Benarroch Jedlicki, J. (2026). A Unifying Lens on Reward Uncertainty in RLHF. arXiv. 预印本 第 35 章
  66. Handa, K., Gal, Y., Pavlick, E., Goodman, N., Andreas, J., Tamkin, A., and Li, B. Z. (2024). Bayesian Preference Elicitation with Language Models. arXiv. 预印本 第 35 章
  67. Hatgis-Kessell, S., Knox, W. B., Booth, S., and Stone, P. (2025). Influencing Humans to Conform to Preference Models for RLHF. Transactions on Machine Learning Research. 第 36 章
  68. Heckel, R., Shah, N. B., Ramchandran, K., and Wainwright, M. J. (2019). Active ranking from pairwise comparisons and when parametric assumptions do not help. The Annals of Statistics. 第 36 章
  69. Hejna, I. D. J., and Sadigh, D. (2022). Few-Shot Preference Learning for Human-in-the-Loop RL. Conference on Robot Learning. 第 36 章
  70. Hejna, J., and Sadigh, D. (2023). Inverse Preference Learning: Preference-based RL without a Reward Function. NeurIPS 2023. 第 36 章
  71. Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D. (2024). Contrastive Preference Learning: Learning from Human Feedback without RL. ICLR 2024. 第 36 章
  72. Herin, M., Perny, P., and Sokolovska, N. (2024). Noise-Tolerant Active Preference Learning for Multicriteria Choice Problems. Algorithmic Decision Theory. 第 36 章
  73. Hong, J., Bhatia, K., and Dragan, A. (2023). On the Sensitivity of Reward Inference to Misspecified Human Models. ICLR 2023. 第 36 章
  74. Hopkins, M., Kane, D., Lovett, S., and Mahajan, G. (2020). Noise-tolerant, Reliable Active Classification with Comparison Queries. Conference on Learning Theory. 第 36 章
  75. Hosseinmardi, H., Ghasemian, A., Rivera-Lanas, M., Horta Ribeiro, M., West, R., and Watts, D. J. (2024). Causally estimating the effect of YouTube’s recommender system using counterfactual bots. Proceedings of the National Academy of Sciences. 第 36 章
  76. Hu, X., Li, J., Zhan, X., Jia, Q.-S., and Zhang, Y.-Q. (2024). Query-Policy Misalignment in Preference-Based Reinforcement Learning. ICLR 2024. 第 36 章
  77. Jannach, D., Manzoor, A., Cai, W., and Chen, L. (2021). A Survey on Conversational Recommender Systems. ACM Computing Surveys. 第 36 章
  78. Ji, K., He, J., and Gu, Q. (2024). Reinforcement Learning from Human Feedback with Active Queries. TMLR. 第 35 章
  79. Joachims, T., Swaminathan, A., and Schnabel, T. (2017). Unbiased Learning-to-Rank with Biased Feedback. WSDM 2017. 第 36 章
  80. Johnston, C. M., Vossler, P., Blessenohl, S., and Vayanos, P. (2023). Deploying a Robust Active Preference Elicitation Algorithm on MTurk: Experiment Design, Interface, and Evaluation for COVID-19 Patient Prioritization. EAAMO 2023. 第 36 章
  81. Kalimeris, D., Bhagat, S., Kalyanaraman, S., and Weinsberg, U. (2021). Preference Amplification in Recommender Systems. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. 第 36 章
  82. Kalinin, S. V., Liu, Y., Biswas, A., Duscher, G., Pratiush, U., Roccapriore, K., Ziatdinov, M., and Vasudevan, R. (2024). Human-in-the-loop: The future of Machine Learning in Automated Electron Microscopy. Microscopy Today. doi:10.1093/mictod/qaad096. 第 36 章
  83. Kane, D. M., Lovett, S., Moran, S., and Zhang, J. (2017). Active classification with comparison queries. FOCS 2017. 第 36 章
  84. Karlekar, S., Zheng, C., Saebo, M., Beltran-Velez, N., Yu, S., Bowlan, J., Kucer, M., and Blei, D. (2026). Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences. ICLR 2026 RSI Workshop. 研讨会论文 第 35 章
  85. Karwowski, J., Hayman, O., Bai, X., Kiendlhofer, K., Griffin, C., and Skalse, J. (2024). Goodhart's Law in Reinforcement Learning. ICLR 2024. 第 36 章
  86. Katkuri, S., Kawada, M., and Wachs, J. (2026). Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning. arXiv. 预印本 第 35 章
  87. Kayal, A., Vakili, S., Toni, L., Shiu, D.-S., and Bernacchia, A. (2025). Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds. International Conference on Machine Learning. 第 35 章
  88. Kim, G., and Kim, E. (2026). Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback. ICLR 2026. 第 35 章
  89. Kirk, H. R., Leqi, L., Zeng, F., Davidson, H., Vidgen, B., Summerfield, C., and Hale, S. A. (2026). PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users. arXiv. 预印本 第 35 章
  90. Kleine Buening, T., Gan, J., Mandal, D., and Kwiatkowska, M. (2025). Strategyproof Reinforcement Learning from Human Feedback. NeurIPS 2025. 第 36 章
  91. Knox, W. B., Hatgis-Kessell, S., Booth, S., Niekum, S., Stone, P., and Allievi, A. (2024). Models of human preference for learning reward functions. TMLR 2024. 第 36 章
  92. Kobalczyk, K., Astorga, N., Liu, T., and van der Schaar, M. (2025). Active Task Disambiguation with LLMs. ICLR 2025. 第 35 章
  93. Kobalczyk, K., Lin, Z. J., Letham, B., Zhao, Z., Balandat, M., and Bakshy, E. (2026). LILO: Bayesian Optimization with Natural Language Feedback. ICML 2026. 第 35 章
  94. Korbak, T., Perez, E., and Buckley, C. L. (2022). RL with KL penalties is better viewed as Bayesian inference. Findings of the Association for Computational Linguistics: EMNLP 2022. doi:10.18653/v1/2022.findings-emnlp.77. 第 35 章
  95. Kristiadi, A., Strieth-Kalthoff, F., Skreta, M., Poupart, P., Aspuru-Guzik, A., and Pleiss, G. (2024a). A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules? International Conference on Machine Learning. 第 35 章
  96. Kuric, E., Demcak, P., and Krajcovic, M. (2026). Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design? arXiv. 预印本 第 35 章
  97. Kutulakos, Z., and Slade, P. (2024). Simulating human-in-the-loop optimization of exoskeleton assistance to compare optimization algorithm performance. bioRxiv. 预印本 第 36 章
  98. Kveton, B., Li, X., McAuley, J., Rossi, R., Shang, J., Wu, J., and Yu, T. (2025). Active Learning for Direct Preference Optimization. arXiv. 预印本 第 35 章
  99. Kwa, T., Thomas, D., and Garriga-Alonso, A. (2024). Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification. NeurIPS 2024. 第 36 章
  100. Laidlaw, C., Bronstein, E., Guo, T., Feng, D., Berglund, L., Svegliato, J., Russell, S., and Dragan, A. (2025). AssistanceZero: Scalably Solving Assistance Games. ICML 2025. 第 36 章
  101. Lang, L., Foote, D., Russell, S., Dragan, A., Jenner, E., and Emmons, S. (2024). When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback. NeurIPS 2024. 第 36 章
  102. Lazzaro, J., Buffelli, D., Shiu, D.-s., and Vakili, S. (2026). A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback. International Conference on Artificial Intelligence and Statistics. 第 35 章
  103. Lee, K., Smith, L., Dragan, A., and Abbeel, P. (2021a). B-Pref: Benchmarking Preference-Based Reinforcement Learning. NeurIPS 2021 Datasets and Benchmarks. 第 36 章
  104. Lee, K., Smith, L., and Abbeel, P. (2021b). PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training. ICML 2021. 第 36 章
  105. Lee, U. H., Shetty, V. S., Franks, P. W., Tan, J., Evangelopoulos, G., Ha, S., and Rouse, E. J. (2023). User preference optimization for control of ankle exoskeletons using sample efficient active learning. Science Robotics. 第 36 章
  106. Li, X., Zhao, H., and Gu, Q. (2024b). Feel-Good Thompson Sampling for Contextual Dueling Bandits. International Conference on Machine Learning. 第 35 章
  107. Li, B. Z., Tamkin, A., Goodman, N., and Andreas, J. (2025b). Eliciting Human Preferences with Language Models. ICLR 2025. 第 35 章
  108. Li, W., Oh, C., and Li, S. (2026d). General Exploratory Bonus for Optimistic Exploration in RLHF. International Conference on Learning Representations. 第 35 章
  109. Li, Z., Liao, Y.-C., and Holz, C. (2026f). Preference-Guided Prompt Optimization for Text-to-Image Generation. CHI 2026. 第 35 章
  110. Liao, Y.-C., Belo, J., Moon, H.-S., Steimle, J., and Feit, A. M. (2026). Efficient Human-in-the-Loop Optimization via Priors Learned from User Models. CHI 2026. 第 35 章
  111. Lin, Z. J., Astudillo, R., Frazier, P., and Bakshy, E. (2022). Preference Exploration for Efficient Bayesian Optimization with Multiple Outcomes. International Conference on Artificial Intelligence and Statistics. 第 35 章
  112. Lin, Y., Seto, S., ter Hoeve, M., Metcalf, K., Theobald, B.-J., Wang, X., … Zhang, T. (2024a). On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization. Findings of the Association for Computational Linguistics: EMNLP 2024. doi:10.18653/v1/2024.findings-emnlp.940. 第 35 章
  113. Lin, X., Dai, Z., Verma, A., Ng, S.-K., Jaillet, P., and Low, B. K. H. (2024b). Prompt Optimization with Human Feedback. ICML 2024 MHFAIA Workshop (no formal proceedings). 研讨会论文 第 35 章
  114. Lindner, D., Turchetta, M., Tschiatschek, S., Ciosek, K., and Krause, A. (2021). Information Directed Reward Learning for Reinforcement Learning. NeurIPS 2021. 第 36 章
  115. Liu, T., Astorga, N., Seedat, N., and van der Schaar, M. (2024a). Large Language Models to Enhance Bayesian Optimization. International Conference on Learning Representations. 第 35 章
  116. Liu, Z., Chen, C., Du, C., Lee, W. S., and Lin, M. (2024b). Sample-Efficient Alignment for LLMs. NeurIPS 2024 LanGame Workshop. 研讨会论文 第 35 章
  117. Liu, T., Qin, Z., Wu, J., Shen, J., Khalman, M., Joshi, R., … Wang, X. (2025a). LiPO: Listwise Preference Optimization through Learning-to-Rank. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). doi:10.18653/v1/2025.naacl-long.121. 第 35 章
  118. Liu, N., Hu, X. E., Savas, Y., Baum, M. A., Berinsky, A. J., Chaney, A. J. B., … Stewart, B. M. (2025b). Short-term exposure to filter-bubble recommendation systems has limited polarization effects: Naturalistic experiments on YouTube. Proceedings of the National Academy of Sciences. 第 36 章
  119. Liu, N., Sun, C., Klinkner, K., and Malmasi, S. (2026a). Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph. arXiv. 预印本 第 35 章
  120. Liu, C., Ling, S., and Jacobson, A. (2026b). GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization. ACM Transactions on Graphics. doi:10.1145/3811293. 第 35 章
  121. Lou, X., Yan, D., Shen, W., Yan, Y., Xie, J., and Zhang, J. (2024). Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown. arXiv (withdrawn from ICLR 2025). 预印本 第 35 章
  122. Ma, Q., Gao, D., Cai, R., Zhao, B., Zhou, H., Zhang, J., and Zhao, Z. (2026). Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization. COLM 2026. 第 35 章
  123. Mahmud, S., Nakamura, M., and Zilberstein, S. (2025). MAPLE: A Framework for Active Preference Learning Guided by Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v39i26.34964. 第 35 章
  124. Mansoury, M., Abdollahpouri, H., Pechenizkiy, M., Mobasher, B., and Burke, R. (2020). Feedback Loop and Bias Amplification in Recommender Systems. CIKM 2020. 第 36 章
  125. Martin, R. M., and Collins, S. H. (2026). Improving CMA-ES Convergence Speed, Efficiency, and Reliability in Noisy Robot Optimization Problems. Evolutionary Computation. 第 36 章
  126. Martin, C., Boutilier, C., Meshi, O., and Sandholm, T. (2024). Model-Free Preference Elicitation. Thirty-Third International Joint Conference on Artificial Intelligence. 第 36 章
  127. McElfresh, D. C., Chan, L., Doyle, K., Sinnott-Armstrong, W., Conitzer, V., Schaich Borg, J., and Dickerson, J. P. (2021). Indecision Modeling. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v35i7.16746. 第 36 章
  128. Mehta, V., Belakaria, S., Das, V., Neopane, O., Dai, Y., Bogunovic, I., … Neiswanger, W. (2025). Sample Efficient Preference Alignment in LLMs via Active Exploration. COLM 2025. 第 35 章
  129. Melikidze, D., Schneider, M., Lam, J., Wertich, M., Hakimi, I., Pásztor, B., and Krause, A. (2026). ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning. ICML 2026. 第 35 章
  130. Melo, L. C., Tigas, P., Abate, A., and Gal, Y. (2024). Deep Bayesian Active Learning for Preference Modeling in Large Language Models. NeurIPS 2024. 第 35 章
  131. Menn, J., Kober, M., Brunzema, P., Stenger, D., and Trimpe, S. (2026a). Local Preferential Bayesian Optimization. arXiv. 预印本 第 35 章
  132. Meta Platforms, Inc. (2026h). BoTorch PairwiseGP source code pairwise_gp.py. GitHub. 软件 第 35 章
  133. Meta Platforms, Inc. (2026l). CHANGELOG (versions 1.2 to 1.3). GitHub. 软件 第 35 章
  134. Mikkola, P., Todorović, M., Järvi, J., Rinke, P., and Kaski, S. (2020). Projective Preferential Bayesian Optimization. International Conference on Machine Learning. 第 36 章
  135. Muldrew, W., Hayes, P., Zhang, M., and Barber, D. (2024). Active Preference Learning for Large Language Models. ICML 2024. 第 35 章
  136. Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., … Piot, B. (2024). Nash Learning from Human Feedback. ICML 2024. 第 35 章
  137. Nan, T., Li, X., Kroer, C., and Lin, T. (2026). Efficient Exploration for Iterative Nash Preference Optimization. arXiv. 预印本 第 35 章
  138. Nguyen, S., Liu, X., and Senanayake, R. (2026). CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM. International Conference on Machine Learning (ICML 2026). 第 35 章
  139. Nisan, N., and Segal, I. (2006). The communication requirements of efficient allocations and supporting prices. Journal of Economic Theory. 第 36 章
  140. Niwa, R., Yoshida, S., Koyama, Y., and Ushiku, Y. (2025). Cooperative Design Optimization through Natural Language Interaction. UIST 2025. 第 35 章
  141. Noothigattu, R., Peters, D., and Procaccia, A. (2020). Axioms for Learning from Pairwise Comparisons. Advances in Neural Information Processing Systems. 第 36 章
  142. Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J. (2020). Dueling Posterior Sampling for Preference-Based Reinforcement Learning. Conference on Uncertainty in Artificial Intelligence. 第 36 章
  143. Oh, G., Lee, J., Park, J., Yu, Y., Bae, W., and Noh, J. (2026b). Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs. ICLR 2026 Workshop: I Can't Believe It's Not Better (ICBINB). 研讨会论文 第 35 章
  144. Oko, K., Ulichney, A., Haghtalab, N., and Bao, H. (2026). Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner. International Conference on Machine Learning. 第 35 章
  145. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems. 第 35 章
  146. Pan, A., Bhatia, K., and Steinhardt, J. (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. ICLR 2022. 第 36 章
  147. Panickssery, A., Bowman, S. R., and Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS 2024. 第 35 章
  148. Papadimitriou, C. H., and Tsitsiklis, J. N. (1987). The Complexity of Markov Decision Processes. Mathematics of Operations Research. 第 36 章
  149. Pásztor, B., Kassraie, P., and Krause, A. (2024). Bandits with Preference Feedback: A Stackelberg Game Perspective. Advances in Neural Information Processing Systems. doi:10.52202/079017-0383. 第 35 章
  150. Peng, Y.-H., Bigham, J. P., and Wu, J. (2026). Efficient Personalization of Generative User Interfaces. arXiv. 预印本 第 35 章
  151. Piriyakulkij, W. T., Kuleshov, V., and Ellis, K. (2023). Active Preference Inference using Language Models and Probabilistic Reasoning. NeurIPS 2023 FMDM Workshop. 研讨会论文 第 35 章
  152. Poddar, S., Wan, Y., Ivison, H., Gupta, A., and Jaques, N. (2024). Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning. Advances in Neural Information Processing Systems 37 (NeurIPS 2024). doi:10.52202/079017-1664. 第 35 章
  153. Pratiush, U., Roccapriore, K. M., Liu, Y., Duscher, G., Ziatdinov, M., and Kalinin, S. V. (2025). Building Workflows for Interactive Human in the Loop Automated Experiment (hAE) in STEM-EELS. Digital Discovery. doi:10.1039/d5dd00033e. 第 36 章
  154. Pukdee, R., Balcan, M.-F., and Ravikumar, P. (2026). What Does Preference Learning Recover from Pairwise Comparison Data? ICML 2026. 第 36 章
  155. Qiu, L., Sha, F., Allen, K., Kim, Y., Linzen, T., and van Steenkiste, S. (2026). Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models. Nature Communications. doi:10.1038/s41467-025-67998-6. 第 35 章
  156. Qu, Z., Zhang, M., Kong, M., Li, X., Shang, Z., Wang, Z., … Dai, Z. (2026). T-POP: Test-Time Personalization with Online Preference Feedback. International Conference on Machine Learning. 第 35 章
  157. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. 第 35 章
  158. Rafailov, R., Chittepu, Y., Park, R., Sikchi, H., Hejna, J., Knox, B., Finn, C., and Niekum, S. (2024). Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms. NeurIPS 2024. 第 36 章
  159. Rajagopalan, R., Dutta, D., Wei, Y.-L., and Roy Choudhury, R. (2026). Personalized Image Generation via Human-in-the-loop Bayesian Optimization. International Conference on Machine Learning. 第 35 章
  160. Ramos, M. C., Michtavy, S. S., Porosoff, M. D., and White, A. D. (2026). Bayesian Optimization of Catalysis With In-Context Learning. ACS Central Science. doi:10.1021/acscentsci.5c02418. 第 35 章
  161. Ranković, B., Griffiths, R.-R., and Schwaller, P. (2026). Large language models as uncertainty-calibrated optimizers for experimental discovery. Nature Machine Intelligence. doi:10.1038/s42256-026-01283-z. 第 35 章
  162. Rodrigues, C., Vas, O., DCosta, I. A., and Prabhakaran, N. K. (2026). When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model. arXiv. 预印本 第 35 章
  163. Rychert, A., Spagnolo, G., and Posashkov, E. (2025). Reproducibility Study of Large Language Model Bayesian Optimization. arXiv. 预印本 第 35 章
  164. Saha, A., Pacchiano, A., and Lee, J. (2023). Dueling RL: Reinforcement Learning with Trajectory Preferences. International Conference on Artificial Intelligence and Statistics. 第 36 章
  165. Saracay, I., Schmidt, L., and Guestrin, C. (2026). Beyond expert users: agents should help users construct preferences, not just elicit them. Conference on Language Modeling (COLM 2026). 第 35 章
  166. Scheid, A., Boursier, E., Durmus, A., Jordan, M. I., Ménard, P., Moulines, E., and Valko, M. (2024). Optimal Design for Reward Modeling in RLHF. arXiv. 预印本 第 35 章
  167. Schoinas, E., Rastogi, A., Carter, A., Granley, J., and Beyeler, M. (2025). Evaluating Deep Human-in-the-Loop Optimization for Retinal Implants Using Sighted Participants. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). doi:10.1109/embc58623.2025.11253762. 第 35 章
  168. Seshadri, P., Cahyawijaya, S., Odumakinde, A., Singh, S., and Goldfarb-Tarrant, S. (2026). Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2026.acl-long.2192. 第 35 章
  169. Sfikas, K., Liapis, A., and Yannakakis, G. N. (2023). Controllable Exploration of a Design Space via Interactive Quality Diversity. arXiv (parts published at GECCO 2023). 预印本 第 36 章
  170. Shah, N. B., Balakrishnan, S., Bradley, J., Parekh, A., Ramchandran, K., and Wainwright, M. J. (2016). Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence. Journal of Machine Learning Research. 第 36 章
  171. Shao, K., Wang, J., Pei, X., and Mesbah, A. (2026). Adaptive KappaSharp: Condition-Number Shaping for Preferential Bayesian Optimization. arXiv. 预印本 第 36 章
  172. Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., … Perez, E. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. 第 36 章
  173. Shen, Y., Sun, H., and Ton, J.-F. (2025a). Active Reward Modeling: Adaptive Preference Labeling for Large Language Model Alignment. International Conference on Machine Learning. 第 35 章
  174. Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., and Vosoughi, S. (2025). Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics. 第 35 章
  175. Shields, B. J., Stevens, J., Li, J., Parasram, M., Damani, F., Alvarado, J. I. M., … Doyle, A. G. (2021). Bayesian reaction optimization as a tool for chemical synthesis. Nature. 第 36 章
  176. Singh, U., Chakraborty, S., Suttle, W. A., Sadler, B. M., Asher, D. E., Sahu, A. K., … Bedi, A. S. (2024). Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach. International Conference on Learning Representations (ICLR 2026). 第 36 章
  177. Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024). Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. ICLR 2024. 第 35 章
  178. Skalse, J., Howe, N. H. R., Krasheninnikov, D., and Krueger, D. (2022). Defining and Characterizing Reward Hacking. Advances in Neural Information Processing Systems 35 (NeurIPS 2022). 第 36 章
  179. Skalse, J., Farrugia-Roberts, M., Russell, S., Abate, A., and Gleave, A. (2023). Invariance in Policy Optimisation and Partial Identifiability in Reward Learning. ICML 2023. 第 36 章
  180. Smith, J. E., and Winkler, R. L. (2006). The Optimizer’s Curse: Skepticism and Postdecision Surprise in Decision Analysis. Management Science. 第 36 章
  181. Song, Y., Swamy, G., Singh, A., Bagnell, J. A., and Sun, W. (2024). The Importance of Online Data: Understanding Preference Fine-tuning via Coverage. NeurIPS 2024. 第 35 章
  182. Sun, H., Shen, Y., and Ton, J.-F. (2025). Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives. ICLR 2025. 第 35 章
  183. Surana, R., Li, X., Yu, S., Shen, Y. J., Wang, C., Yu, T., … Wu, J. (2026). MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization. arXiv. 预印本 第 35 章
  184. Takagi, H. (2001). Interactive evolutionary computation: fusion of the capabilities of EC optimization and human evaluation. Proceedings of the IEEE. 第 36 章
  185. Takagi, H., and Pallez, D. (2009). Paired Comparisons-based Interactive Differential Evolution. NaBIC 2009. 第 36 章
  186. Takeno, S., Nomura, M., and Karasuyama, M. (2023). Towards Practical Preferential Bayesian Optimization with Skew Gaussian Processes. International Conference on Machine Learning. 第 36 章
  187. Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., … Piot, B. (2024). Generalized Preference Optimization: A Unified Approach to Offline Alignment. International Conference on Machine Learning. 第 35 章
  188. Thies, S. M. A. R., Bengs, V., Kaufmann, T., Vollmer, S. J., and Hüllermeier, E. (2026a). Calibrated Preference Learning: The Case of Label Ranking. International Conference on Machine Learning (ICML 2026). 第 35 章 第 36 章
  189. Thies, S. M. A. R., Alfaro, J. C., and Bengs, V. (2026b). MORE-PLR: multi-output regression employed for partial label ranking. Machine Learning 115. 第 36 章
  190. Vayanos, P., Ye, Y., McElfresh, D., Dickerson, J., and Rice, E. (2020). Robust Active Preference Elicitation. arXiv (journal version not found). 预印本 第 36 章
  191. Vendrov, I., Lu, T., Huang, Q., and Boutilier, C. (2020). Gradient-based Optimization for Bayesian Preference Elicitation. AAAI 2020. 第 36 章
  192. Verma, A., Dai, Z., Lin, X., Jaillet, P., and Low, B. K. H. (2025). Neural Dueling Bandits: Preference-Based Optimization with Human Feedback. International Conference on Learning Representations. 第 35 章
  193. Viappiani, P., and Boutilier, C. (2010). Optimal Bayesian Recommendation Sets and Myopically Optimal Choice Query Sets. Advances in Neural Information Processing Systems. 第 36 章
  194. Viappiani, P., and Boutilier, C. (2020). On the equivalence of optimal recommendation sets and myopically optimal query sets. Artificial Intelligence. 第 36 章
  195. Wang, T., and Boutilier, C. (2003). Incremental Utility Elicitation with the Minimax Regret Decision Criterion. Proceedings of the Eighteenth International Joint Conference on Artificial Intelligence (IJCAI-03). 第 36 章
  196. Wang, Y., and Pei, Y. (2024). A comprehensive survey on interactive evolutionary computation in the first two decades of the 21st century. Applied Soft Computing. doi:10.1016/j.asoc.2024.111950. 第 36 章
  197. Wang, Y., Liu, Q., and Jin, C. (2023a). Is RLHF More Difficult than Standard RL? NeurIPS 2023. 第 35 章 第 36 章
  198. Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., … Sui, Z. (2024a). Large Language Models are not Fair Evaluators. ACL 2024. 第 35 章 第 36 章
  199. Wang, Y., Sun, Z., Zhang, J., Xian, Z., Biyik, E., Held, D., and Erickson, Z. (2024b). RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback. International Conference on Machine Learning. 第 35 章
  200. Wang, H., Branke, J., and Poloczek, M. (2025a). Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble. Advances in Neural Information Processing Systems 38. doi:10.52202/085713-4124. 第 36 章
  201. Wen, J., Zhong, R., Khan, A., Perez, E., Steinhardt, J., Huang, M., … Feng, S. (2025). Language Models Learn to Mislead Humans via RLHF. ICLR 2025. 第 36 章
  202. Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., and Dragan, A. (2025). On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback. ICLR 2025. 第 36 章
  203. Won, Y., Lee, H., Hwang, H., and Seo, M. (2025). Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization. arXiv. 预印本 第 35 章
  204. Wu, R., and Sun, W. (2024). Making RL with Preference-based Feedback Efficient via Randomization. ICLR 2024. 第 35 章
  205. Wu, Y., Verma, S., Lee, J., Xiong, F., Zhang, P., Awadelkarim, A., … Hill, S. (2026). LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization. Findings of the Association for Computational Linguistics: ACL 2026. doi:10.18653/v1/2026.findings-acl.490. 第 35 章
  206. Xia, F., Liu, H., Yue, Y., and Li, T. (2025). Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents. Findings of the Association for Computational Linguistics: ACL 2025. doi:10.18653/v1/2025.findings-acl.519. 第 35 章
  207. Xie, T., Foster, D. J., Krishnamurthy, A., Rosset, C., Awadallah, A., and Rakhlin, A. (2025). Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF. ICLR 2025. 第 35 章
  208. Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T. (2024). Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint. ICML 2024. 第 35 章
  209. Xu, W., Wang, W., Jiang, Y., Svetozarevic, B., and Jones, C. (2024b). Principled Preferential Bayesian Optimization. International Conference on Machine Learning. 第 35 章
  210. Xu, Y., Ruis, L., Rocktäschel, T., and Kirk, R. (2025a). Investigating Non-Transitivity in LLM-as-a-Judge. International Conference on Machine Learning. 第 35 章
  211. Yang, A. X., Robeyns, M., Coste, T., Shi, Z., Wang, J., Bou-Ammar, H., and Aitchison, L. (2024). Bayesian Reward Models for LLM Alignment. ICLR 2024 SeT LLM Workshop; ICML 2024 SPIGM Workshop. 研讨会论文 第 35 章
  212. Yang, J., Hu, Z., Qiu, C., Deng, Z., Jiao, X., and Zhou, T. (2026a). Quantifying and Mitigating Self-Preference Bias of LLM Judges. arXiv. 预印本 第 35 章
  213. Yang, D., Stante, S., Redhardt, F., Libon, L., Kassraie, P., Hakimi, I., Pásztor, B., and Krause, A. (2026b). RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models. EurIPS 2025 EIML Workshop. 研讨会论文 第 35 章
  214. Yuan, Y., Hao, J., Ma, Y., Dong, Z., Liang, H., Liu, J., … Zheng, Y. (2024). Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback. ICLR 2024. 第 36 章
  215. Yuan, X., Chen, Z., Zhang, J., Xiong, H., Ye, N., Li, Y., and Gu, Q. (2026). Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery. ICLR 2026. 第 35 章
  216. Zhang, Z.-Y., Han, S., Yao, H., Niu, G., and Sugiyama, M. (2024a). Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought. International Conference on Machine Learning. 第 35 章
  217. Zhang, S., Yu, D., Sharma, H., Zhong, H., Liu, Z., Yang, Z., … Wang, Z. (2024b). Self-Exploring Language Models: Active Preference Elicitation for Online Alignment. TMLR. 第 35 章
  218. Zhang, X., Huang, D., Kaski, S., and Martinelli, J. (2025a). PABBO: Preferential Amortized Black-Box Optimization. ICLR 2025. 第 35 章
  219. Zhao, H., Ye, C., Gu, Q., and Zhang, T. (2025). Sharp Analysis for KL-Regularized Contextual Bandits and RLHF. NeurIPS 2025. 第 35 章
  220. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., … Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems. doi:10.52202/075280-2020. 第 35 章
  221. Zhu, B., Jordan, M., and Jiao, J. (2023). Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. International Conference on Machine Learning. 第 35 章
  222. Zhu, L., Huang, X., and Sang, J. (2024). How Reliable is Your Simulator? Analysis on the Limitations of Current LLM-based User Simulators for Conversational Recommendation. Companion Proceedings of the ACM Web Conference 2024. doi:10.1145/3589335.3651955. 研讨会论文 第 35 章
  223. Zhuang, S., and Hadfield-Menell, D. (2020). Consequences of Misaligned AI. NeurIPS 2020. 第 36 章
  224. Zintgraf, L. M., Roijers, D. M., Linders, S., Jonker, C. M., and Nowé, A. (2018). Ordered Preference Elicitation Strategies for Supporting Multi-Objective Decision Making. AAMAS 2018. 第 36 章