揭示深度线性网络中随机梯度下降的噪声如何反映特征学习进程。
Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
- 将 SGD 动态建模为各向异性状态依赖噪声的朗之万方程。
- 发现每种特征的扩散峰值先于其完全学习,且噪声影响分布形态。
- 适用于研究深度网络训练动态的理论分析者,尤其关注噪声作用者。
深度线性网络(DLNs)是解析理解深度神经网络训练动态的可解析模型。尽管已知梯度下降在 DLNs 中呈现鞍点到鞍点的动力学,但随机梯度下降(SGD)噪声在此过程中的影响仍不清晰。本文研究了在鞍点到鞍点阶段训练 DLNs 时的 SGD 动态。我们将其建模为具有各向异性、状态依赖噪声的随机朗之万动力学。在权重对齐且平衡的假设下,导出动力学的精确分解,得到一组沿每个模式的一维随机微分方程。结果表明,模式上的最大扩散发生在该特征被完全学习之前。我们还推导出每个模式的稳态分布:无标签噪声时,其边缘分布与梯度流一致;有标签噪声时,则近似为玻尔兹曼分布。实验验证表明,这些理论结果在权重未对齐或不平衡时仍定性成立。结果表明,SGD 噪声编码了特征学习进展的信息,但并未根本改变鞍点到鞍点的动力学结构。
原文摘要 · Abstract (English)
Deep linear networks (DLNs) are used as an analytically tractable model of the training dynamics of deep neural networks. While gradient descent in DLNs is known to exhibit saddle-to-saddle dynamics, the impact of stochastic gradient descent (SGD) noise on this regime remains poorly understood. We investigate the dynamics of SGD during training of DLNs in the saddle-to-saddle regime. We model the training dynamics as stochastic Langevin dynamics with anisotropic, state-dependent noise. Under the assumption of aligned and balanced weights, we derive an exact decomposition of the dynamics into a system of one-dimensional per-mode stochastic differential equations. This establishes that the maximal diffusion along a mode precedes the corresponding feature being completely learned. We also derive the stationary distribution of SGD for each mode: in the absence of label noise, its marginal distribution along specific features coincides with the stationary distribution of gradient flow, while in the presence of label noise it approximates a Boltzmann distribution. Finally, we confirm experimentally that the theoretical results hold qualitatively even without aligned or balanced weights. These results establish that SGD noise encodes information about the progression of feature learning but does not fundamentally alter the saddle-to-saddle dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。