推理微调的泛化能力受优化、数据和模型能力共同影响,非简单记忆。
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
- 通过长思维链监督,发现泛化呈先降后升的动态过程
- 高质量长思维链数据可带来跨领域性能提升,弱模型仅模仿表面文本
- 强模型能内化可迁移的推理模式,但可能损害安全性
大语言模型后训练中普遍认为监督微调(SFT)会记忆而强化学习(RL)更擅长泛化。本文针对具有长思维链(CoT)监督的推理SFT重新审视该观点,发现跨领域泛化并非缺失,而是由优化动态、训练数据与基础模型能力共同决定的条件性现象。部分报告的失败实为欠优化导致:跨域性能先下降后回升,长期训练才能显现泛化优势。数据质量与结构至关重要:低质量解法普遍削弱泛化,而经验证的长思维链轨迹则带来稳定跨域收益。模型能力不可或缺:强模型即使从简单的算术游戏也能内化可迁移的程序性模式(如回溯),弱模型仅模仿表层冗余表达。这种泛化具有不对称性:推理能力提升的同时安全性能下降,促使我们重新思考‘推理SFT能否泛化’的问题,转为关注其在何种条件下以及以何种代价实现泛化。
原文摘要 · Abstract (English)
A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit this claim for reasoning SFT with long chain-of-thought (CoT) supervision and find that cross-domain generalization is not absent but conditional, jointly shaped by optimization dynamics, training data, and base-model capability. Some reported failures are under-optimization artifacts: cross-domain performance first degrades before recovering and improving with extended training (a dip-and-recovery pattern), so shorttraining checkpoints can underestimate generalization. Data quality and structure both matter: low-quality solutions broadly hurt generalization,while verified long-CoT traces yield consistent cross-domain gains. Model capability is essential: stronger models internalize transferable procedural patterns (e.g., backtracking) even from a toy arithmetic game, while weaker ones imitate surface verbosity. This generalization is asymmetric, however: reasoning improves while safety degrades, reframing the question from whether reasoning SFT generalizes to under what conditions and at what cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。