arXiv:2605.28389cs.CL2026-05

让大模型一边解题一边自检,训练更快更准。

FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning

  • 解题与自检合并为一步生成,减少训练时间。
  • 新方法使奖励持续提升,小模型效果提高23%。
  • 适合想高效训练数学推理模型的研究者。

尽管大语言模型在数学推理方面取得显著进展,但其仍难以准确判断自身解答的正确性。现有自检方法通常将解题与验证分为两个独立任务,导致训练时间大幅增加。本文提出FABSVer,将解题与自检融合为单次生成过程,显著降低训练开销,并联合优化两者能力。我们进一步发现:随着训练推进,奖励趋于平缓,根源在于策略受固定参考模型限制。为此提出动态参考模型更新(DRMU),打破奖励天花板,实现持续奖励增长。在多个数学基准测试中,FABSVer在三种模型规模下均实现更优的自检与推理性能,训练时间仅为现有方法的51%–71%。分析揭示模型存在不同的学习阶段,且随模型规模增大,验证与答案奖励差距明显缩小。

原文摘要 · Abstract (English)

While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own solutions. Existing approaches that equip models with self-verification typically treat solution generation and verification as two separate tasks, leading to substantially increased training time. In this paper, we propose FABSVer, which fuses these two tasks into a single generation pass, dramatically reducing training overhead while jointly optimizing both capabilities. We further identify a convergence bottleneck both theoretically and empirically: as training progresses, the reward reaches a plateau because the policy is constrained by a fixed reference model. To overcome this, we introduce Dynamic Reference Model Update (DRMU), which raises the reward ceiling and enables sustained reward growth. Extensive experiments on math benchmarks demonstrate that FABSVer achieves superior self-verification and reasoning performance across three model scales, while requiring only 51%--71% of the training time of existing methods. Analysis further reveals distinct learning phases in how models acquire self-verification, and that the gap between verify and answer rewards shrinks noticeably as model size increases.

数学推理自验证训练加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。