教机器听懂人类反馈,让大模型更符合人类意图。
Reinforcement Learning from Human Feedback

- 基于人类反馈强化学习,分步优化语言模型对齐性。
- 涵盖指令微调、奖励建模到直接对齐的完整训练流程。
- 适合想系统掌握大模型后训练技术的研究者与工程师。
从人类反馈中进行强化学习(RLHF)已成为构建大规模机器学习系统的关键工具。本文全面介绍后训练模型的核心方法,面向具备一定量化背景的读者,围绕标准的RLHF流程展开。内容包括RLHF的起源、发展里程碑及其在强化学习背景下的基础原理;详细阐述从指令微调、奖励模型训练,到拒绝采样、强化学习、在线蒸馏和直接对齐算法的每一步优化过程。同时探讨了RLHF在经济学、哲学与最优控制等领域的交叉渊源。结尾讨论合成数据、工具使用、角色训练与评估等前沿问题及开放挑战。配套提供代码库、模型对比工具与教学课程,构成后训练语言模型学习的一站式资源。
原文摘要 · Abstract (English)
Reinforcement learning from human feedback (RLHF) has become a crucial tool to build the latest machine learning systems at scale. The field grew around the core methods of RLHF into today's broader suite of post-training techniques. In this book, we give a comprehensive introduction to the core methods for post-training models for people with some level of quantitative background, organized around the canonical RLHF recipe. The book starts with what RLHF does and why it was created, with seminal technical milestones in its young history and a primer on reinforcement learning context needed to understand the book. The core of the book details every optimization stage in using RLHF, from starting with instruction tuning to training a reward model and finally all of rejection sampling, reinforcement learning, on-policy distillation, and direct alignment algorithms. The book also discusses broader topics, such as the origins of RLHF -- both in recent literature and in a convergence of disparate fields of science in economics, philosophy, and optimal control. The book concludes with advanced topics -- understudied or emerging research questions in synthetic data, tool-use, character training, and evaluation -- and open questions for the field. The book is released with a variety of companion resources, including a codebase, a library to compare model completions from within post-training stages, and an educational course, to be a one-stop shop for learning all foundational concepts for post-training language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。