arXiv:2605.17967cs.AI2026-05

揭示大模型微调效果不稳的根源:早期去噪有效,后期过拟合反而有害。

Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective

论文配图:Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective
图 1 · 摘自论文原文
  • 从词元交互角度分析微调过程,发现主要清除噪声交互
  • 微调早期仅短暂去噪,后续持续训练会引入过拟合交互
  • 为大模型训练提供早停指导,适合关注微调策略的研究者

本文探讨监督微调(SFT)在小规模深度神经网络中广泛有效,却在大语言模型(LLMs)上表现不一致甚至有害这一科学问题。基于交互解释的新进展表明,词元间的交互能忠实量化LLM的推理模式。研究发现,SFT过程中交互结构的演变可有效解释其对LLMs效果不一的现象:(1) SFT主要消除类噪声交互,极少获得可靠新交互;(2) 该去噪阶段极短,此后持续微调易引入过拟合交互。研究在多个LLM和数据集上验证了上述发现,为早停策略提供了新见解,并给出大模型训练的实践指导。

原文摘要 · Abstract (English)

This paper explores a scientific question in supervised fine-tuning (SFT): why SFT is broadly effective for small-scale deep neural networks, yet can produce inconsistent or even detrimental effects when applied to large language models (LLMs). Recent advances in interaction-based explanations suggest that interactions between words/tokens provide a faithful metric for quantifying the inference patterns encoded by LLMs. We find that the evolution of interactions during SFT can effectively explain the inconsistent effectiveness of SFT for LLMs. Specifically, we find that (1) SFT primarily removes noise-like interactions, while rarely acquiring reliable new interactions. (2) This denoising stage is extremely brief, after which continued fine-tuning tends to introduce overfitted interactions. We validate these findings across multiple LLMs and datasets. Our findings provide new insights into early stopping and offer practical guidance for LLM training.

大模型微调交互分析早停策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。