arXiv:2505.24864cs.CLcs.AI2025-05NeurIPS被引 164

通过长期强化学习,让大模型发现全新推理策略。

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

  • 采用KL控制与策略重置,实现持续强化学习训练。
  • 在多种任务上超越基础模型,即使多次尝试仍失败的场景也能解决。
  • 适合研究长时序强化学习与模型推理能力扩展的学者。

近期以推理为中心的语言模型进展表明,强化学习(RL)是使模型对可验证奖励进行对齐的有前景方法。然而,仍有争议:RL是否真正拓展了模型的推理能力,还是仅放大了基模型分布中已存在的高奖励输出?持续增加RL计算量是否能稳定提升推理性能?本文挑战现有假设,证明长期强化学习(ProRL)可挖掘出基模型无法触及的新型推理策略,即使在大量采样下也无效。我们提出ProRL,结合KL散度控制、参考策略重置和多样化任务集。实证分析显示,经RL训练的模型在广泛pass@k评估中持续优于基模型,包括基模型无论尝试多少次都完全失败的情况。进一步表明,推理边界提升与基模型任务能力及训练时长强相关,说明RL能随时间探索并填充解空间的新区域。这些发现为理解RL如何实质性拓展语言模型推理边界提供了新视角,并为未来长时序推理强化学习研究奠定基础。模型权重已公开:https://huggingface.co/nvidia/Nemotron-Research-Reasoning-Qwen-1.5B

原文摘要 · Abstract (English)

Recent advances in reasoning-centric language models have highlighted reinforcement learning (RL) as a promising method for aligning models with verifiable rewards. However, it remains contentious whether RL truly expands a model's reasoning capabilities or merely amplifies high-reward outputs already latent in the base model's distribution, and whether continually scaling up RL compute reliably leads to improved reasoning performance. In this work, we challenge prevailing assumptions by demonstrating that prolonged RL (ProRL) training can uncover novel reasoning strategies that are inaccessible to base models, even under extensive sampling. We introduce ProRL, a novel training methodology that incorporates KL divergence control, reference policy resetting, and a diverse suite of tasks. Our empirical analysis reveals that RL-trained models consistently outperform base models across a wide range of pass@k evaluations, including scenarios where base models fail entirely regardless of the number of attempts. We further show that reasoning boundary improvements correlates strongly with task competence of base model and training duration, suggesting that RL can explore and populate new regions of solution space over time. These findings offer new insights into the conditions under which RL meaningfully expands reasoning boundaries in language models and establish a foundation for future work on long-horizon RL for reasoning. We release model weights to support further research: https://huggingface.co/nvidia/Nemotron-Research-Reasoning-Qwen-1.5B

强化学习推理能力大模型长期训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。