arXiv:2501.08425cs.LGmath.AP2025-01

从偏微分方程视角解析SGD的优化机制与收敛行为。

Is Stochastic Gradient Descent Effective? A PDE Perspective on Machine Learning processes

  • 用福克-普朗克型抛物方程建模SGD,揭示权重演化规律。
  • 发现初始阶段权重聚集于最近局部极小值,后期依赖随机波动逃逸。
  • 首次给出非凸、退化扩散下的收敛性证明,适用于真实神经网络训练。

本文从抛物型福克-普朗克方程的角度分析了广泛应用于监督学习中的随机梯度下降(SGD)方法,其目标是通过最小化非凸损失函数来优化神经网络权重。尽管福克-普朗克方程已有长期研究基础,但当势能函数非凸或扩散矩阵退化时,相关理论几乎空白,这正是本研究的核心挑战。我们识别出两个不同阶段:初始阶段为“漂移阶段”,损失函数驱动权重向最近的局部极小值集中,本文给出了该集中的定量估计;随后进入“扩散阶段”,随机扰动帮助模型逃离次优局部极小。我们分析了平均离开时间(MET),并证明了其上下界。最后,针对非凸代价函数和退化扩散矩阵下的渐近收敛问题,传统方法失效,我们采用对偶与熵方法,提出新理论结果。这些成果深化了随机优化与偏微分方程之间的联系,回答了机器学习中若干基本问题:SGD逃离劣质极小值需要多久?神经网络参数在SGD下是否收敛?训练初期参数如何演化?

原文摘要 · Abstract (English)

In this paper we analyze the behaviour of the stochastic gradient descent (SGD), a widely used method in supervised learning for optimizing neural network weights via a minimization of non-convex loss functions. Since the pioneering work of E, Li and Tai (2017), the underlying structure of such processes can be understood via parabolic PDEs of Fokker-Planck type, which are at the core of our analysis. Even if Fokker-Planck equations have a long history and a extensive literature, almost nothing is known when the potential is non-convex or when the diffusion matrix is degenerate, and this is the main difficulty that we face in our analysis. We identify two different regimes: in the initial phase of SGD, the loss function drives the weights to concentrate around the nearest local minimum. We refer to this phase as the drift regime and we provide quantitative estimates on this concentration phenomenon. Next, we introduce the diffusion regime, where stochastic fluctuations help the learning process to escape suboptimal local minima. We analyze the Mean Exit Time (MET) and prove upper and lower bounds of the MET. Finally, we address the asymptotic convergence of SGD, for a non-convex cost function and a degenerate diffusion matrix, that do not allow to use the standard approaches, and require new techniques. For this purpose, we exploit two different methods: duality and entropy methods. We provide new results about the dynamics and effectiveness of SGD, offering a deep connection between stochastic optimization and PDE theory, and some answers and insights to basic questions in the Machine Learning processes: How long does SGD take to escape from a bad minimum? Do neural network parameters converge using SGD? How do parameters evolve in the first stage of training with SGD?

SGD分析偏微分方程非凸优化收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。