揭示深度网络权重谱的动态演化规律,解析异常值与主成分的协同变化机制。
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer

- 构建双层平均场理论,同时追踪主成分与异常值的谱演化
- 发现μP参数化下异常值增长稳定,且可跨宽度迁移学习率
- 适用于小输出任务;大输出任务需重构谱主成分结构
我们研究了在(随机)梯度下降训练下,宽深度神经网络中隐藏权重谱的演化。提出一种两级动力学平均场理论(DMFT),能联合追踪具有非独立方向的稀疏特征集中的主成分与异常值动态。该框架应用于两种情形:(1) 无限宽非线性网络在均值场/μP缩放下;(2) 深度线性网络在比例高维极限中,其宽度、输入维度与样本量以固定比率发散。理论预测了异常值随训练时间、网络宽度、输出尺度和初始化方差的演化规律。在深度线性网络中,μP参数化实现宽度一致的异常值动态与超参数转移,包括主导的NTK模式向稳定性边缘(EoS)的稳定增长。相比之下,尽管NTK参数化最终收敛到稳定的宽网络极限,其异常值动态仍强烈依赖于宽度。我们证明,这种主成分+异常值的图景可描述小输出通道的简单任务,但涉及大量输出的任务(如ImageNet分类或GPT语言建模)则更符合谱主成分的重构。我们构建了一个具有广泛输出通道的简化模型重现该现象,并表明在足够宽的网络中,谱边缘依然收敛。
原文摘要 · Abstract (English)
We study the evolution of hidden-weight spectra in wide neural networks trained by (stochastic) gradient descent. We develop a two-level dynamical mean-field theory (DMFT) that jointly tracks bulk and outlier spectral dynamics for spiked ensembles whose spike directions remain statistically dependent on the random bulk. We apply this framework to two settings: (1) infinite-width nonlinear networks in mean-field/$μ$P scaling and (2) deep linear networks in the proportional high-dimensional limit, where width, input dimension, and sample size diverge with fixed ratios. Our theory predicts how outliers evolve with training time, width, output scale, and initialization variance. In deep linear networks, $μ$P yields width-consistent outlier dynamics and hyperparameter transfer, including width-stable growth of the leading NTK mode toward the edge of stability (EoS). In contrast, NTK parameterization exhibits strongly width-dependent outlier dynamics, despite converging to a stable large-width limit. We show that this bulk+outlier picture is descriptive of simple tasks with small output channels, but that tasks involving large numbers of outputs (ImageNet classification or GPT language modeling) are better described by a restructuring of the spectral bulk. We develop a toy model with extensive output channels that recapitulates this phenomenon and show that edge of the spectrum still converges for sufficiently wide networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。