arXiv:2507.12709cs.LG2025-07被引 8

揭示了SGD训练中权重谱演化的数学规律。

From SGD to Spectra: A Theory of Neural Network Weight Dynamics

  • 用连续时间随机微分方程建模权重动态,连接微观优化与宏观谱变化。
  • 发现奇异值平方遵循具有排斥效应的迪森布朗运动,稳定分布呈幂律尾部。
  • 首次理论解释网络权重谱的'主体+尾部'结构,适合研究深度学习机理者。

深度神经网络革新了机器学习,但其训练动态仍缺乏理论解释。本文建立了一个连续时间、矩阵值的随机微分方程(SDE)框架,严格连接了SGD的微观动态与权重矩阵奇异值谱的宏观演化。我们推导出精确的SDE,表明平方奇异值遵循具有特征排斥效应的迪森布朗运动,并刻画了其稳态分布为具有幂律尾部的伽马型密度,首次从理论上解释了训练后网络中普遍观察到的‘主体+尾部’谱结构。通过在Transformer和MLP架构上的受控实验,我们验证了理论预测,展示了基于SDE的预测与实际谱演化间具有定量一致性,为理解深度学习为何有效提供了严格基础。

原文摘要 · Abstract (English)

Deep neural networks have revolutionized machine learning, yet their training dynamics remain theoretically unclear-we develop a continuous-time, matrix-valued stochastic differential equation (SDE) framework that rigorously connects the microscopic dynamics of SGD to the macroscopic evolution of singular-value spectra in weight matrices. We derive exact SDEs showing that squared singular values follow Dyson Brownian motion with eigenvalue repulsion, and characterize stationary distributions as gamma-type densities with power-law tails, providing the first theoretical explanation for the empirically observed 'bulk+tail' spectral structure in trained networks. Through controlled experiments on transformer and MLP architectures, we validate our theoretical predictions and demonstrate quantitative agreement between SDE-based forecasts and observed spectral evolution, providing a rigorous foundation for understanding why deep learning works.

神经网络权重动态谱分析随机微分方程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。