arXiv:2601.20829cs.LGcs.AI2026-01被引 2

让大模型从错误路径中学习,突破推理训练的瓶颈。

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning

  • 用罕见错误推理前缀引导模型,激发饱和问题中的潜在学习信号。
  • 在标准强化学习停滞处仍能提升性能,效果相当于新收集中等难度题。
  • 适合想提升模型鲁棒性、解决训练卡顿的研究者和工程师。

随着可验证奖励的强化学习(RLVR)显著提升大语言模型的推理能力,新的瓶颈出现:越来越多的训练问题进入饱和状态,即模型几乎每次推演都能正确回答。此类问题中,奖励提供的学习信号微弱。虽然收集更难的问题是自然应对方式,但成本高且越来越困难。本文提出失败前缀条件化方法,通过聚焦于罕见错误轨迹的前缀,引导探索向易出错的推理状态转移,从而增强模型从误导性早期推理中恢复的能力。实验表明,该方法在标准RLVR停滞处持续提升性能,增益相当于训练于新收集的中等难度问题。进一步分析显示,该方法降低了误导性错误前缀下的性能下降,尽管对正确早期推理的遵循略有削弱。最后,我们证明迭代式刷新失败前缀可带来性能平台期后的额外提升。结果表明,饱和问题仍蕴含有效学习信号,而失败前缀条件化是一种高效挖掘手段。

原文摘要 · Abstract (English)

As Reinforcement Learning with Verifiable Rewards (RLVR) substantially improves the reasoning abilities of large language models (LLMs), a new bottleneck emerges: more training problems become saturated, that is, the LLM answers the questions correctly for nearly every rollout. On such problems, rewards provide little useful learning signal. While collecting harder problems is a natural response, it is costly and increasingly difficult. We propose failure-prefix conditioning, a simple method that unlocks the remaining signal in saturated problems by shifting exploration toward failure-prone reasoning states. By conditioning on prefixes of rare incorrect trajectories, the method improves the model's ability to recover from misleading early reasoning. We observe that failure-prefix conditioning consistently improves performance where standard RLVR stalls, and achieves gains comparable to training on newly collected medium-difficulty problems. We further analyze the model's robustness, finding that our method reduces performance degradation under misleading failure prefixes, albeit with a mild trade-off in adherence to correct early reasoning. Finally, we demonstrate that an iterative approach, which refreshes failure prefixes during training, unlocks additional gains after performance plateaus. Overall, our results show that saturated problems still contain valuable learning signal, and that failure-prefix conditioning provides an effective way to unlock it.

推理训练强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。