arXiv:2602.08489cs.LGcs.CL2026-02被引 2

让大模型推理过程更稳健,提升通用性和训练效率

Beyond Correctness: Learning Robust Reasoning via Transfer

  • 用可迁移奖励机制评估推理片段的实用性
  • 在MATH500上准确率提升3.6个百分点,训练步数减少2.5倍
  • 适合关注推理稳定性与高效训练的研究者

强化学习结合可验证奖励(RLVR)虽提升了大模型的推理能力,但仅关注最终答案正确性,忽视推理过程的鲁棒性。本文提出基于可迁移奖励的强化学习(RLTR),认为鲁棒推理应能跨越模型边界被复用,通过测试一个模型的推理前缀能否指导另一模型得出正确答案来实现。该方法促使模型生成稳定、可解释且真正通用的推理过程。实验显示,RLTR在保持更高采样一致性的同时提升最终答案准确率,在MATH500上较RLVR提升3.6个百分点(Maj@64),且仅需约2.5倍少的训练步数即可达到相当平均准确率,兼具更强推理可靠性与显著更高的样本效率。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has recently strengthened LLM reasoning, but its focus on final answer correctness leaves a critical gap: it does not ensure the robustness of the reasoning process itself. We adopt a simple philosophical view, robust reasoning should remain useful beyond the mind that produced it, and treat reasoning as a form of meaning transfer that must survive truncation, reinterpretation, and continuation. Building on this principle, we introduce Reinforcement Learning with Transferable Reward (RLTR), which operationalizes robustness via transfer reward that tests whether a partial reasoning prefix from one model can guide a separate model to the correct answer. This encourages LLMs to produce reasoning that is stable, interpretable, and genuinely generalizable. Our approach improves sampling consistency while improving final answer accuracy, and it reaches comparable performance in substantially fewer training steps. For example, on MATH500, RLTR achieves a +3.6%p gain in Maj@64 compared to RLVR and matches RLVR's average accuracy with roughly 2.5x fewer training steps, providing both more reliable reasoning and significantly more sample efficient.

大模型推理强化学习鲁棒性训练效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。