arXiv:2609.07755cs.LGmath.OC2026-09

解析神经网络训练中泛化能力的动态变化,揭示权重衰减下梯度下降的收敛机制。

A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay

  • 基于输入数据划分空间,分解总体误差为数据、优化与预测波动三部分。
  • 提出局部近似均匀性,给出各层预测波动的显式上下界。
  • 解释了泛化提升的必要条件与延迟泛化现象,关联‘悟道’现象的理论机制。

理解泛化仍是机器学习的核心挑战,需同时考虑数据、模型结构与训练动态。本文建立理论框架,刻画这些因素如何共同影响训练全过程中的泛化性能。研究在ℓ²损失下由梯度下降(GD)配合权重衰减训练的一类广泛神经网络,证明了GD收敛至经验损失全局最小值的邻域。通过依据输入数据划分空间,将总体误差分解为数据误差、优化误差和预测波动误差,并分别进行上界估计。尤其针对衡量学习函数振荡的预测波动误差,提出(局部)近似均匀性假设,推导出沿训练轨迹的逐单元与逐层显式边界。这些边界带来两个重要启示:一是改进泛化的必要条件,解释不同层泛化行为差异;二是充分条件描述延迟泛化,提供对‘悟道’(grokking)现象的理论刻画。

原文摘要 · Abstract (English)

Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training dynamics. In this paper, we develop a theoretical framework that characterizes how these factors jointly shape generalization performance throughout training. More precisely, we study a broad class of neural networks trained under the $\ell^2$ loss by gradient descent (GD) with weight decay, and prove the convergence of GD to a neighbourhood of the global minimizers of the empirical loss. By partitioning the space based on the input data, we then decompose the population error into data error, optimization error, and prediction variation error, and bound them separately. In particular, for the prediction variation error, which measures the oscillations of the learned function, we propose (local) approximate homogeneity and derive explicit cellwise and layerwise bounds for its evolution along the training trajectory. These bounds yield two important implications: a necessary condition of improved generalization explains differences in layerwise generalization behavior; a sufficient condition describes delayed generalization and provides a theoretical characterization of grokking.

泛化分析梯度下降权重衰减神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。