让模型记住定理而非答案,提升数学推理泛化能力
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning

- 用定理应用代替答案记忆,引导模型关注推理规则
- MATH数据集上提升8.8%,GeoQA上提升20.27%
- 仅微调MLP层即可达全层效果,凸显前馈网络关键作用
监督微调(SFT)虽广泛用于任务适配,但近期研究发现其系统性损害推理泛化能力。我们指出问题根源并非记忆本身,而是记忆目标:标准SFT促使模型利用并记忆问题-答案对中的表面伪相关性,导致对输入表层变化敏感。为此,我们提出定理微调(Theorem-SFT),将监督重点转向显式定理应用,教会模型如何调用规则而非记忆答案样式。该方法在多个基准和模型族中均带来稳定提升:在MATH(LLaMA3.2-3B-Instruct)上提升8.8%,在GeoQA(Qwen2.5-VL-7B-Instruct)上提升20.27%,且无需模态特异性重训练。仅微调MLP层即可达到全层性能,表明前馈组件是推理规则的主要承载位置。研究重新定义了争论焦点:泛化失败并非源于记忆机制,而是记错了归纳目标。
原文摘要 · Abstract (English)
Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization. We argue the root cause is not memorization itself, but its target: vanilla SFT drives models to exploit and memorize spurious surface correlations in problem-solution pairs, leaving them brittle to superficial input variations. To address this, we propose Theorem-SFT, which reorients supervision toward explicit theorem application by teaching models how rules are invoked rather than what answers look like. Theorem-SFT yields consistent gains across benchmarks and model families: +8.8% on MATH (LLaMA3.2-3B-Instruct) and +20.27% on GeoQA (Qwen2.5-VL-7B-Instruct) without modality-specific re-training. Fine-tuning MLP layers alone matches full-layers performance, implicating feed-forward components as the primary locus of reasoning rules. Our findings reframe the debate: Generalization failures stem not from memorization as a mechanism, but from memorizing the wrong inductive targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。