arXiv:2510.21245cs.LGmath.OC2025-10

分析了梯度下降在懒惰训练下的收敛性,揭示其稳定性和快速优化机制。

Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime

  • 基于连续时间随机微分方程建模梯度下降过程
  • 证明算法在有限时间与宽度下可指数收敛到最优解
  • 适用于关注优化稳定性与理论保障的研究者

连续时间模型为深度学习优化算法的训练动态提供了重要洞察。本文在懒惰训练范式下,对随机梯度朗之万动力学(SGLD)进行非渐近收敛分析。SGLD是随机梯度下降的伊藤型随机微分方程(SDE)连续时间近似。在损失函数海森矩阵满足正则性条件下,我们证明:(i) SGLD在训练过程中以高概率保持非退化核;(ii) 其期望值以指数速度收敛至经验风险最小化器,并给出了优化差距的有限时间与有限宽度界。数值实验在回归设置中验证了理论结果。

原文摘要 · Abstract (English)

Continuous-time models provide important insights into the training dynamics of optimization algorithms in deep learning. In this work, we establish a non-asymptotic convergence analysis of stochastic gradient Langevin dynamics (SGLD), which is an Itô stochastic differential equation (SDE) approximation of stochastic gradient descent in continuous time, in the lazy training regime. We show that, under regularity conditions on the Hessian of the loss function, SGLD with multiplicative and state-dependent noise (i) yields a non-degenerate kernel throughout the training process with high probability, and (ii) achieves exponential convergence to the empirical risk minimizer in expectation, and we establish finite-time and finite-width bounds on the optimality gap. We corroborate our theoretical findings with numerical examples in the regression setting.

优化理论随机梯度收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。