通过优化推理轨迹,提升大模型数学推理的初始能力。
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
- 用双目标优化强化弱信心正确数据、抑制强信心错误数据
- 在6个基准上平均提升6分Pass@1和5分Pass@k
- 适合需要长期推理优化的研究者与开发者
通过鼓励自我探索,基于可验证奖励的强化学习(RLVR)显著提升了大语言模型的数学推理能力。作为RLVR的起点,监督微调(SFT)对新思维链轨迹的记忆能力,为后续探索提供了关键初始化。然而,现有研究主要关注RLVR训练中的探索促进,而对探索感知的SFT研究不足。为此,我们提出离线探索感知(OXA)微调。OXA优化两个目标:促进低置信度但已验证的教师蒸馏数据,以内化先前未捕捉的推理模式;抑制高置信度错误的自蒸馏数据,将错误模式的概率质量重新分配至潜在正确候选。在6个基准上的实验结果表明,OXA持续提升数学推理性能,尤其在Qwen2.5-1.5B-Math上,平均相比传统SFT提升+6 Pass@1和+5 Pass@k。关键的是,OXA提高了初始策略熵,且性能增益在整个长周期RLVR训练中保持,证明了其长期价值。
原文摘要 · Abstract (English)
Through encouraging self-exploration, reinforcement learning from verifiable rewards (RLVR) has significantly advanced the mathematical reasoning capabilities of large language models. As the starting point for RLVR, the capacity of supervised fine-tuning (SFT) to memorize new chain-of-thought trajectories provides a crucial initialization that shapes the subsequent exploration landscape. However, existing research primarily focuses on facilitating exploration during RLVR training, leaving exploration-aware SFT under-explored. To bridge this gap, we propose Offline eXploration-Aware (OXA) fine-tuning. Specifically, OXA optimizes two objectives: promoting low-confidence verified teacher-distillation data to internalize previously uncaptured reasoning patterns, and suppressing high-confidence incorrect self-distillation data to redistribute probability mass of incorrect patterns toward potentially correct candidates. Experimental results across 6 benchmarks show that OXA consistently improves mathematical reasoning performance, especially achieving an average gain of $+6$ Pass@1 and $+5$ Pass@$k$ points compared to conventional SFT on the Qwen2.5-1.5B-Math. Crucially, OXA elevates initial policy entropy, and performance gains persist throughout extensive RLVR training, demonstrating the long-term value of OXA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。