首次精确描述有限宽度网络梯度下降动态,突破传统无限宽理论局限。
Precise gradient descent training dynamics for finite-width multi-layer neural networks
- 提出非渐近状态演化理论,刻画权重分布随迭代变化规律。
- 在有限宽度下实现训练与泛化误差的精确分析,支持非高斯特征。
- 适用于一般多层网络,可指导早停与超参调优,适合理论研究者。
本文首次在样本量与特征维数成比例增长、网络宽度与深度有界的有限宽度比例尺度下,对通用多层神经网络的梯度下降迭代提供了精确的分布表征。我们的非渐近状态演化理论捕捉了第一层权重的高斯波动和深层权重的集中现象,且对非高斯特征仍成立。该理论与现有的神经正切核(NTK)、平均场(MF)理论及张量程序(TP)存在关键差异:其一,理论运行于有限宽度而非无限宽;其二,允许权重从任意初始值演化,超越懒惰训练区域,而NTK与MF对初始化几乎不变或仅弱敏感,TP依赖特殊初始化;其三,可分析一般多层网络的训练与泛化误差,超出统一收敛范式,现有理论大多仅限于两层情形。作为统计应用,我们证明了普通梯度下降可被扩展以在每一步迭代中一致估计泛化误差,用于引导早停与超参数调优。进一步理论表明,即使模型误设,梯度下降所学模型仍保持单索引函数结构,有效信号为真实信号与初始化的线性组合。
原文摘要 · Abstract (English)
In this paper, we provide the first precise distributional characterization of gradient descent iterates for general multi-layer neural networks under the canonical single-index regression model, in the `finite-width proportional regime' where the sample size and feature dimension grow proportionally while the network width and depth remain bounded. Our non-asymptotic state evolution theory captures Gaussian fluctuations in first-layer weights and concentration in deeper-layer weights, and remains valid for non-Gaussian features. Our theory differs from existing neural tangent kernel (NTK), mean-field (MF) theories and tensor program (TP) in several key aspects. First, our theory operates in the finite-width regime whereas these existing theories are fundamentally infinite-width. Second, our theory allows weights to evolve from individual initializations beyond the lazy training regime, whereas NTK and MF are either frozen at or only weakly sensitive to initialization, and TP relies on special initialization schemes. Third, our theory characterizes both training and generalization errors for general multi-layer neural networks beyond the uniform convergence regime, whereas existing theories study generalization almost exclusively in two-layer settings. As a statistical application, we show that vanilla gradient descent can be augmented to yield consistent estimates of the generalization error at each iteration, which can be used to guide early stopping and hyperparameter tuning. As a further theoretical implication, we show that despite model misspecification, the model learned by gradient descent retains the structure of a single-index function with an effective signal determined by a linear combination of the true signal and the initialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。