数学推理强的模型未必通用,强化学习比监督微调更利于能力迁移。
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
- 对比强化学习与监督微调,发现前者更利于跨领域泛化。
- 20多个模型测试显示,数学优者在其他任务上表现普遍不佳。
- 适合关注模型泛化能力与训练方法优化的研究者。
数学推理已成为大语言模型进步的代表,新模型在MATH和AIME等基准上迅速超越人类水平。然而,随着数学排行榜周周更新,有必要追问:这些提升是反映更广泛的问题解决能力,还是仅限于狭窄的过拟合?为此,我们评估了超过20个开源权重的推理微调模型,涵盖数学、科学问答、智能体规划、编程及标准指令遵循等任务。出人意料的是,大多数在数学上表现优异的模型无法将其优势迁移到其他领域。为深入研究此现象,我们在Qwen3-14B模型上进行受控实验,仅使用数学数据但采用不同微调方法。结果表明,强化学习(RL)微调的模型在各领域间具有良好泛化能力,而监督微调(SFT)模型常丧失通用能力。潜在空间表示与词元分布偏移分析显示,SFT导致显著的表征与输出漂移,而RL则保留了通用领域结构。研究提示需重新思考标准后训练流程,尤其是对基于SFT蒸馏数据推进推理模型的依赖。
原文摘要 · Abstract (English)
Math reasoning has become the poster child of progress in large language models (LLMs), with new models rapidly surpassing human-level performance on benchmarks like MATH and AIME. But as math leaderboards improve week by week, it is worth asking: do these gains reflect broader problem-solving ability or just narrow overfitting? To answer this question, we evaluate over 20 open-weight reasoning-tuned models across a broad suite of tasks, including math, scientific QA, agent planning, coding, and standard instruction-following. We surprisingly find that most models that succeed in math fail to transfer their gains to other domains. To rigorously study this phenomenon, we conduct controlled experiments on Qwen3-14B models using math-only data but different tuning methods. We find that reinforcement learning (RL)-tuned models generalize well across domains, while supervised fine-tuning (SFT)-tuned models often forget general capabilities. Latent-space representation and token-space distribution shift analyses reveal that SFT induces substantial representation and output drift, while RL preserves general-domain structure. Our results suggest a need to rethink standard post-training recipes, particularly the reliance on SFT-distilled data for advancing reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。