优化数学推理大模型后训练,发现先SFT再RL效果最佳。
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
- 提出塑性天花板框架,拆解性能上限为SFT基础与RL提升空间。
- 实验证明先SFT后RL优于同步训练,避免过早收敛与不稳。
- 建议在SFT稳定或轻微过拟合时转入RL,数据量决定潜力上限。
监督微调(SFT)和强化学习(RL)主导数学推理大模型的后训练,但对专家轨迹的依赖方式不同。为探索最优利用策略,本文提出塑性天花板框架,通过分解最终性能上限为基础SFT表现与后续RL可塑性(即最大提升空间),实证分析了后训练格局。大规模基准测试表明,顺序式SFT-then-RL流程优于同步方法,克服了稳定性差与过早收敛问题。进一步得出精确缩放指南:(1)在SFT的稳定或轻度过拟合阶段转入RL,可构建稳健基础并保留充足提升空间;(2)反驳“少即是多”假设,证实数据规模决定后训练潜力,轨迹难度则作为性能乘数;(3)SFT的最小验证损失是选择专家轨迹以最大化最终性能的关键指标。研究提供可操作的专家轨迹利用指导。
原文摘要 · Abstract (English)
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) dominate the post-training landscape for mathematical reasoning, yet differ fundamentally in their reliance on expert trajectories. To understand the optimal way to harness these trajectories for maximizing performance, we propose the Plasticity-Ceiling Framework. This framework empirically grounds the post-training landscape by decomposing the final performance ceiling into the foundational SFT performance and the subsequent RL plasticity (i.e., the maximum improvement via RL). Through extensive benchmarking, we establish the Sequential SFT-then-RL pipeline as the superior standard, overcoming the stability and premature convergence deficits inherent in synchronized approaches. Furthermore, we derive precise scaling guidelines: (1) Transitioning to RL at the Stable or Mild Overfitting Regime of SFT maximizes the final ceiling by securing a robust SFT foundation with substantial RL plasticity; (2) Refuting the ``Less is More'' hypothesis in SFT-then-RL scaling, we demonstrate that Data Scale determines the primary post-training potential, while Trajectory Difficulty acts as a performance multiplier; and (3) The Minimum Validation Loss of SFT serves as a reliable indicator for selecting the expert trajectories that maximize the ultimate performance ceiling. Our findings provide actionable guidelines for extracting maximum value from expert trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。