RLVR让大模型在答题时更注重正确推理过程。
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
- 用可验证奖励强化学习,引导模型优化推理步骤。
- 数学与编程任务中,推理成功率显著提升,Pass@K明显改善。
- 适合研究大模型推理能力或强化学习应用的学者。
近期基于深度思考(CoT)的进展,特别是深思-1(DeepSeek-R1)采用的组相对策略优化算法,引发了对大语言模型(LLMs)中可验证奖励强化学习(RLVR)潜力的广泛关注。尽管RLVR有望通过自由探索提升推理能力,但其是否真正增强推理而非仅提高采样效率仍存争议。本文系统研究了RLVR对LLM推理的影响。我们重新审视Pass@K实验,证明RLVR能扩展数学与编程任务的推理边界。为此,我们提出新评估指标CoT-Pass@K,同时考量最终答案与中间推理步骤。此外,我们构建理论框架解释RLVR的激励机制,表明即使奖励仅基于答案正确性,也能促进正确推理。训练动态分析显示,模型早期即被激励进行正确推理,经大量评估验证,推理质量显著提升。这些发现有力支持了RLVR增强LLM推理的潜力,为理解其机制与性能改进提供了重要洞见。
原文摘要 · Abstract (English)
Recent advancements in long chain-of-thought (CoT) reasoning, particularly through the Group Relative Policy Optimization algorithm used by DeepSeek-R1, have led to significant interest in the potential of Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs). While RLVR promises to improve reasoning by allowing models to learn from free exploration, there remains debate over whether it truly enhances reasoning abilities or simply boosts sampling efficiency. This paper systematically investigates the impact of RLVR on LLM reasoning. We revisit Pass@K experiments and demonstrate that RLVR can extend the reasoning boundary for both mathematical and coding tasks. This is supported by our introduction of a novel evaluation metric, CoT-Pass@K, which captures reasoning success by accounting for both the final answer and intermediate reasoning steps. Furthermore, we present a theoretical framework explaining RLVR's incentive mechanism, demonstrating how it can encourage correct reasoning even when rewards are based solely on answer correctness. Our analysis of RLVR's training dynamics reveals that it incentivizes correct reasoning early in the process, with substantial improvements in reasoning quality confirmed through extensive evaluations. These findings provide strong evidence of RLVR's potential to enhance LLM reasoning, offering valuable insights into its mechanisms and performance improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。