arXiv:2602.24110cs.AIcs.CL2026-02被引 2

通过细粒度修正错误步骤,挽救近似正确的推理路径以扩展探索空间。

Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance

  • 识别错误首步并逐步修正部分正确轨迹,实现精准反馈。
  • 探索多样性提升13.5%,数学推理平均准确率达46.6%。
  • 适合需要稳定探索与高鲁棒性的复杂推理任务研究者。

强化学习从可验证奖励(RLVR)已成为提升大推理模型复杂推理能力的有效范式。然而,基于结果的监督存在根本缺陷:对大部分正确但因少量失误而失败的轨迹,惩罚程度与完全错误轨迹相同。这种粗粒度反馈导致模型丢弃宝贵的近似正确轨迹,造成轨迹多样性下降,过早压缩探索空间。过程奖励模型在测试时缩放中展现出可靠的步骤级验证能力,但将其信号直接作为密集奖励使用效果不佳。现有方法尝试引入离策略全轨迹替换,但常超出策略分布且未能有效利用模型自身生成的近似正确轨迹,无法缓解探索空间收缩问题。为此,我们提出SCOPE(步骤级纠正以维持在线探索),利用过程奖励模型定位低质量轨迹中的首个错误步骤,并实施细粒度、步骤级的离策略修正。通过对部分正确轨迹进行精确优化,本方法有效挽救近似正确路径,在实验中将多样性提升13.5%,显著维持广阔探索空间。大量实验证明,该方法达到新最优性能,在数学推理任务上平均准确率达46.6%,在分布外推理任务上仍保持53.4%的鲁棒泛化能力。

原文摘要 · Abstract (English)

Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing the complex reasoning capabilities of Large Reasoning Models. However, standard outcome-based supervision suffers from a critical limitation that penalizes trajectories that are largely correct but fail due to several missteps as heavily as completely erroneous ones. This coarse feedback signal causes the model to discard valuable largely correct rollouts, leading to a degradation in rollout diversity that prematurely narrows the exploration space. Process Reward Models have demonstrated efficacy in providing reliable step-wise verification for test-time scaling, naively integrating these signals into RLVR as dense rewards proves ineffective.Prior methods attempt to introduce off-policy guided whole-trajectory replacement that often outside the policy model's distribution, but still fail to utilize the largely correct rollouts generated by the model itself and thus do not effectively mitigate the narrowing of the exploration space. To address these issues, we propose SCOPE (Step-wise Correction for On-Policy Exploration), a novel framework that utilizes Process Reward Models to pinpoint the first erroneous step in suboptimal rollouts and applies fine-grained, step-wise off-policy rectification. By applying precise refinement on partially correct rollout, our method effectively salvages partially correct trajectories and increases diversity score by 13.5%, thereby sustaining a broad exploration space. Extensive experiments demonstrate that our approach establishes new state-of-the-art results, achieving an average accuracy of 46.6% on math reasoning and exhibiting robust generalization with 53.4% accuracy on out-of-distribution reasoning tasks.

强化学习推理增强探索多样性过程奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。