解析大模型强化学习后训练的内在机制,揭示成功关键。
Demystifying Reinforcement Learning Post-Training of Language Models

- 用可验证奖励拆解强化学习流程,分析各环节影响
- 发现提示分布决定虚假奖励的影响程度
- 适合想理解强化学习原理的NLP研究者
强化学习(RL)后训练已成为提升大语言模型(LLMs)能力的强大框架,使其在推理、数学和编程方面表现卓越。然而对许多研究者而言,经典强化学习原理仍如黑箱。本文剖析强化学习后训练算法的每一步,通过在受控简化环境中使用可验证奖励,考察基础模型先验分布、奖励信号粒度、提示分布多样性及模型规模如何共同塑造训练结果。我们以策略输出分布熵为视角,比较预训练、监督微调(SFT)与强化学习后训练所学分布,揭示各阶段如何影响模型确定性。研究发现,所谓‘虚假奖励’的效果取决于后训练时使用的提示分布;同时表明强化学习成功与否取决于基础模型是否已对期望行为赋予足够概率质量,这与强化学习中的探索概念密切相关。最终,本文为希望将强化学习纳入工具箱的自然语言处理研究者提供实用指南。
原文摘要 · Abstract (English)
Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities. Yet for many researchers and practitioners, the principles behind classical RL remain a "black box". In this work, we deconstruct the RL post-training algorithm, investigating each step to clarify what is actually happening beneath the surface. By isolating the mechanics of RL with Verifiable Rewards in a controlled and simplified environment, we examine how RL outcomes are shaped by the base model's prior distribution, the granularity of the reward signal, the diversity of the prompt distribution, and model scale. We use the entropy of the policy's output distribution as a lens to compare the distributions learned through pretraining, SFT, and RL post-training, revealing how each stage shapes model certainty. Our investigation sheds light on how these choices interact to affect post-training success. For example, we show that the effect of so-called 'spurious rewards' depends on the prompt distribution used for post-training. We also provide insight into why the success of RL post-training depends on whether the base model already places sufficient probability mass on the desired behavior, linking it to the classical concept of exploration in RL. Ultimately, we provide this primer as a resource to those in the NLP community wishing to incorporate RL as a tool in their toolbox.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。