数学推理能力由预训练中形成的少数关键层决定,后续训练无法改变其重要性。
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
- 通过层级消融实验发现,数学推理依赖少数关键层。
- 移除关键层后数学准确率下降最高达80%。
- 这些层在不同微调方法下保持稳定,适合研究模型可解释性。
大型语言模型在指令微调、强化学习或知识蒸馏后,数学推理能力得到提升。我们探究这些提升是否源于变换器层的重大变化,还是仅对原有结构进行微调。通过在基础模型和训练后变体上进行逐层消融实验,发现数学推理依赖于少数关键层,且在所有后续训练方法中保持重要性。移除这些层会使数学准确率下降高达80%,而事实回忆任务的下降幅度相对较小。这表明,针对数学任务的专用层在预训练阶段已形成,并在后续训练中保持稳定。通过归一化互信息(NMI)度量发现,在这些关键层附近,标记从原始句法聚类向与句法关联较弱但下游任务可能更相关的表示漂移。
原文摘要 · Abstract (English)
Large language models improve at math after instruction tuning, reinforcement learning, or knowledge distillation. We ask whether these gains come from major changes in the transformer layers or from smaller adjustments that keep the original structure. Using layer-wise ablation on base and trained variants, we find that math reasoning depends on a few critical layers, which stay important across all post-training methods. Removing these layers reduces math accuracy by as much as 80%, whereas factual recall tasks only show relatively smaller drops. This suggests that specialized layers for mathematical tasks form during pre-training and remain stable afterward. As measured by Normalized Mutual Information (NMI), we find that near these critical layers, tokens drift from their original syntactic clusters toward representations aligned with tokens less syntactically related but potentially more useful for downstream task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。