揭示深度神经网络中前向后向权重耦合的动态机制,提出可解析求解的简化模型。
Feature Learning Dynamics in Infinite-Depth Neural Networks
- 通过条件高斯分解,分离权重重用带来的前后向耦合项与独立噪声项。
- 训练中耦合项在无限宽极限下仍存在,但随深度增加被抑制,衰减速率为 $O(L^{-2})$。
- 提出神经特征动力学模型,适合研究深层残差网络的特征学习规律。
深度神经网络虽已取得显著成果,但对训练过程中特征演化的机理理解仍不完整,尤其是在大深度极限下。针对深度-μP缩放下的残差网络,已有研究将层索引ℓ视为连续时间tℓ=ℓ/L,得到训练动态的随机微分方程(SDE)描述。关键未解问题是反向传播重复使用前向权重矩阵Wℓ及其转置Wℓ⊤,导致前向特征与反向梯度间存在相关性,其行为和在特征学习中的作用尚不清楚。本文研究在一阶残差网络下权重重用引发的前后向耦合现象。采用条件高斯表示,显式分离由权重重用产生的耦合项与去耦的高斯波动项,且不依赖任何网络极限。初始化时证明该耦合为有限宽度效应,以$O(n^{-1})$速率消失,且对深度一致。然而训练中,SGD引入非平凡的前后向相关项,该结构在无限宽极限下仍保留。关键深度效应在于:在深度-μP缩放下,该保留项为高阶深度项,其跨层累积贡献在$L\to\infty$时趋于零。这一深度抑制机制催生了神经特征动力学(NFD),一种前后向权重解耦的前向-后向SDE系统,保留了训练中生成的特征-梯度协方差结构。在非退化假设下,证明有限网络训练动态收敛至其NFD极限,深度离散误差为$O(L^{-1})$,而原权重重用耦合项衰减更快,为$O(L^{-2})$。这些结果为深度-μP缩放下一阶残差网络的特征学习动态提供了严格的无限深度极限。
原文摘要 · Abstract (English)
Deep neural networks have achieved remarkable success in practice, yet a mechanistic understanding of how features evolve during training remains incomplete, especially in the large-depth limit. For ResNets under depth-$μ$P scaling, prior work treats the layer index $\ell$ as a continuous time $t_\ell = \ell/L$, yielding SDE descriptions of the training dynamics. A key unresolved issue is that backpropagation reuses each forward weight matrix $W_\ell$ through its transpose $W_\ell^\top$, creating correlations between forward features and backward gradients whose behavior and role in feature learning remain unclear. We study this reused-weight forward--backward coupling in one-layer ResNets under depth-$μ$P. Using conditional Gaussian representations, we explicitly separate the coupling terms induced by weight reuse from decoupled Gaussian fluctuations before taking any network limit. At initialization, we prove that the coupling is a finite-width effect and vanishes at rate $O(n^{-1})$, uniformly over depth. During training, however, SGD induces a nontrivial forward--backward correlation term that survives the infinite-width limit. The key depth effect is that, under depth-$μ$P scaling, this surviving term is higher order in depth and its accumulated contribution over layers becomes negligible as $L\to\infty$. This depth-induced suppression motivates Neural Feature Dynamics (NFD), a forward--backward SDE system with decoupled backward weights that retains the feature-gradient covariance structure generated during training. Under nondegeneracy assumptions, we prove that the finite-network training dynamics converge to its NFD limit with an $O(L^{-1})$ depth-discretization error, while the reused-weight coupling term has a faster $O(L^{-2})$ decay. These results provide a rigorous infinite-depth limit for the feature-learning dynamics of one-layer ResNets under depth-$μ$P.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。