arXiv:2602.13942stat.MLcs.LG2026-02被引 1

解释为何大模型微调只需几轮就能见效

A Theoretical Framework for LLM Fine-tuning Using Early Stopping for Non-random Initialization

  • 将神经正切核理论扩展至预训练初始化,建立收敛性保证
  • 发现收敛速度取决于核矩阵特征值衰减速率
  • 可解释多任务微调中的任务向量机制,适合研究者参考

在大语言模型时代,微调预训练模型已成为常态,但其理论基础仍不明确。核心问题在于:为何仅需少量微调轮次即可在多种任务上取得优异表现?本文构建了一个统计框架,结合严格的早停理论与基于注意力的神经正切核(NTK),为微调实践提供新理论洞见。我们形式化地将经典NTK理论 [Jacot et al., 2018] 扩展至非随机(即预训练)初始化,并为基于注意力的微调提供了收敛性保证。关键发现是:收敛速率与由NTK诱导的样本核矩阵的特征值衰减速率密切相关。此外,该框架还可用于解释大模型中多任务任务向量的现象。在真实数据集上的实验验证了理论结果的有效性。

原文摘要 · Abstract (English)

In the era of large language models (LLMs), fine-tuning pretrained models has become ubiquitous. Yet the theoretical underpinning remains an open question. A central question is why only a few epochs of fine-tuning are typically sufficient to achieve strong performance on many different tasks. In this work, we approach this question by developing a statistical framework, combining rigorous early stopping theory with the attention-based Neural Tangent Kernel (NTK) for LLMs, offering new theoretical insights on fine-tuning practices. Specifically, we formally extend classical NTK theory [Jacot et al., 2018] to non-random (i.e., pretrained) initializations and provide a convergence guarantee for attention-based fine-tuning. One key insight provided by the theory is that the convergence rate with respect to sample size is closely linked to the eigenvalue decay rate of the empirical kernel matrix induced by the NTK. We also demonstrate how the framework can be used to explain task vectors for multiple tasks in LLMs. Finally, experiments with modern language models on real-world datasets provide empirical evidence supporting our theoretical insights.

大模型微调神经正切核理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。