arXiv:2504.14945cs.LGcs.AI2025-04NeurIPS被引 280

让大模型通过外部推理数据突破自身能力局限,提升数学推理水平。

Learning to Reason under Off-Policy Guidance

  • 引入离线推理数据混合训练,动态平衡模仿与探索。
  • 在6个数学基准上平均提升超6.4分,分布外任务领先6.2分。
  • 可训练弱模型,适用于难以靠自生成数据进化的场景。

近期大型推理模型(LRMs)的发展表明,通过可验证奖励的强化学习(RLVR)可涌现出多步推理与自我反思等复杂行为。然而,现有RLVR方法本质上为“在线策略”,仅依赖模型自身输出,无法获取超出初始能力的推理能力。为此,我们提出LUFFY(Learning to reason Under off-policy guidance),通过引入离线策略推理轨迹增强RLVR。LUFFY在训练中结合离线演示与在线采样,动态平衡模仿与探索。具体地,采用具有理论收敛保证的Mixed-Policy GRPO框架,并通过正则化重要性采样进行策略塑形,避免混合策略训练中的表面模仿与僵化。相较于以往方法,LUFFY在6个数学基准上平均提升超过+6.4分,在分布外任务中优势达+6.2分以上。最重要的是,我们在原有在线策略RLVR完全失效的场景下,成功训练了弱模型。结果表明,LUFFY突破了在线策略RLVR的根本局限,证明了离线引导在强化学习推理中的巨大潜力。

原文摘要 · Abstract (English)

Recent advances in large reasoning models (LRMs) demonstrate that sophisticated behaviors such as multi-step reasoning and self-reflection can emerge via reinforcement learning with verifiable rewards~(\textit{RLVR}). However, existing \textit{RLVR} approaches are inherently ``on-policy'', limiting learning to a model's own outputs and failing to acquire reasoning abilities beyond its initial capabilities. To address this issue, we introduce \textbf{LUFFY} (\textbf{L}earning to reason \textbf{U}nder o\textbf{FF}-polic\textbf{Y} guidance), a framework that augments \textit{RLVR} with off-policy reasoning traces. LUFFY dynamically balances imitation and exploration by combining off-policy demonstrations with on-policy rollouts during training. Specifically, LUFFY combines the Mixed-Policy GRPO framework, which has a theoretically guaranteed convergence rate, alongside policy shaping via regularized importance sampling to avoid superficial and rigid imitation during mixed-policy training. Compared with previous RLVR methods, LUFFY achieves an over \textbf{+6.4} average gain across six math benchmarks and an advantage of over \textbf{+6.2} points in out-of-distribution tasks. Most significantly, we show that LUFFY successfully trains weak models in scenarios where on-policy RLVR completely fails. These results provide compelling evidence that LUFFY transcends the fundamental limitations of on-policy RLVR and demonstrates the great potential of utilizing off-policy guidance in RLVR.

推理模型强化学习离线引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。