高维下,梯度下降可被精确建模为连续随机过程,揭示收敛机制。
High-dimensional Limit of SGD for Diagonal Linear Networks

- 用随机微分方程描述高维梯度下降,分离梯度噪声与漂移项。
- 推导出确定性偏微分方程,刻画风险、曲率等统计量演化。
- 在合适参数化下,几乎必然指数级收敛至零风险,结果非渐近明确。
理解随机梯度方法的行为是现代机器学习的核心问题。近期研究指出,对角线线性网络是一种简化但富有表现力的设置,可用于分析神经模型的优化与泛化特性。本文表明,在高维情形下,对角线线性网络上的随机梯度下降可被连续动力学良好逼近,该动力学由随机微分方程(SDE)支配,且显式地将漂移与梯度噪声解耦。我们进一步推导出一个确定性偏微分方程,其解能传播迭代器的相关状态,并刻画一类广泛可观测统计量(包括风险、曲率及其他最优性指标)的时间演化。最后,我们在合适参数化下证明,随机动力学全局适定,且以高概率指数级收敛至零风险,从而给出其长时间行为的完全显式非渐近描述。数值模拟验证了理论结果。
原文摘要 · Abstract (English)
Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet expressive setting for analyzing the optimization and generalization properties of neural models. In this work, we show that in the high-dimensional regime, stochastic gradient descent on diagonal linear networks is well-approximated by continuous dynamics governed by a stochastic differential equation (SDE), which explicitly decouples the drift from the gradient noise. We further derive a deterministic partial differential equation whose solution propagates the relevant state of the iterates and characterizes the time evolution of a broad class of observable statistics, including the risk, curvature, and other metrics for optimality. Finally, we show that, under a suitable parametrization, the stochastic dynamics are globally well posed and converge exponentially fast to zero risk with high probability, yielding a fully explicit non-asymptotic description of their long-time behavior. Numerical simulations corroborate our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。