arXiv:2511.04622math.OCcs.LG2025-11被引 9

用微分方程分析Adam优化器,揭示其收敛本质与寻优能力。

ODE approximation for the Adam algorithm: General and overparametrized setting

  • 构建Adam的微分方程模型,刻画其在快慢尺度下的动态行为。
  • 证明收敛点必为Adam向量场的零点,非传统极小点。
  • 在过参数化场景下,可逼近全局最优解,适合深度学习优化分析。

Adam优化器目前是深度学习中最流行的优化方法。本文通过基于常微分方程(ODE)的方法,在快-慢尺度下研究Adam算法。当动量参数固定且步长趋于零时,我们证明Adam算法是特定向量场(称为Adam向量场)流的渐近伪轨迹。利用渐近伪轨迹的性质,建立了Adam算法的收敛性结果。特别地,在一般设定下,若算法收敛,则极限必为Adam向量场的零点,而非目标函数的局部极小或临界点。相反,在过参数化的经验风险最小化设置中,Adam算法能够局部找到全局极小值集合。具体而言,我们证明在全局极小值邻域内,目标函数是该向量场诱导流的李雅普诺夫函数。因此,若Adam算法无限次进入全局极小值邻域,它将收敛至全局极小值集合。

原文摘要 · Abstract (English)

The Adam optimizer is currently presumably the most popular optimization method in deep learning. In this article we develop an ODE based method to study the Adam optimizer in a fast-slow scaling regime. For fixed momentum parameters and vanishing step-sizes, we show that the Adam algorithm is an asymptotic pseudo-trajectory of the flow of a particular vector field, which is referred to as the Adam vector field. Leveraging properties of asymptotic pseudo-trajectories, we establish convergence results for the Adam algorithm. In particular, in a very general setting we show that if the Adam algorithm converges, then the limit must be a zero of the Adam vector field, rather than a local minimizer or critical point of the objective function. In contrast, in the overparametrized empirical risk minimization setting, the Adam algorithm is able to locally find the set of minima. Specifically, we show that in a neighborhood of the global minima, the objective function serves as a Lyapunov function for the flow induced by the Adam vector field. As a consequence, if the Adam algorithm enters a neighborhood of the global minima infinitely often, it converges to the set of global minima.

优化器分析Adam微分方程收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。