纯强化学习训练30亿参数模型,提升推理能力。
Reinforcement Learning is all You Need
- 用计时游戏实现无人类反馈的强化学习训练。
- 在五个基准中四个表现优于基线,泛化能力增强。
- 适合关注纯强化学习推理的科研人员。
受DeepSeek R1通过强化学习实现推理成功启发,我们使用计时游戏对30亿参数语言模型进行纯强化学习训练。该模型在五个基准中的四个上优于基线,表明其超越训练数据的泛化能力。值得注意的是,回答长度与推理质量无关;虽然出现“顿悟时刻”,但并不总能得出正确答案。这些发现凸显了仅用强化学习训练在推理增强方面的潜力,并提示未来需优化奖励机制,以将涌现的洞察与准确性更好关联。
原文摘要 · Abstract (English)
Inspired by the success of DeepSeek R1 in reasoning via reinforcement learning without human feedback, we train a 3B language model using the Countdown Game with pure reinforcement learning. Our model outperforms baselines on four of five benchmarks, demonstrating improved generalization beyond its training data. Notably, response length does not correlate with reasoning quality, and while "aha moments" emerge, they do not always yield correct answers. These findings highlight the potential of RL-only training for reasoning enhancement and suggest future work on refining reward structures to bridge emergent insights with accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。