arXiv:2607.04233math.OCcs.LG2026-07被引 1

统一分析深度网络优化方法收敛性,涵盖Adam等主流算法。

Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks

  • 基于KL不等式统一分析多种梯度优化方法的收敛性。
  • 证明了软激活函数下算法可收敛至临界点,适用于Adam等10余种优化器。
  • 为AI优化算法理论提供新视角,适合研究者与工程师参考。

基于梯度的优化方法是当前人工智能系统中训练深度神经网络(DNN)的首选方法。在实际问题中,通常不使用标准梯度下降(GD),而是采用包含自适应或加速机制的复杂方法,如著名的Adam优化器。本文首次针对具有软激活函数(如softplus、GeLU)的DNN训练,提供了一套通用的统一收敛性分析。该分析适用于包括标准GD、动量法、Nesterov加速梯度(NAG)、RMSprop、Adam、Adamax、Nadam、Nadamax、Adan、AdaBelief、AMSGrad和Yogi在内的多种梯度优化方法。通过引入Kurdyka-Łojasiewicz(KL)不等式理论,证明了这些方法在训练过程中可收敛至临界点。据我们所知,该分析的广度在Adam优化器的特定情况下也是一项新的文献贡献。

原文摘要 · Abstract (English)

Gradient based optimization methods are nowadays the methods of choice for training deep neural networks (DNNs) in artificial intelligence (AI) systems. In practically relevant DNN training problems, one does usually not apply the standard gradient descent (GD) optimization method but instead one employs suitable sophisticated GD optimization methods, which incorporate adaptivity and/or acceleration techniques, such as the famous Adam optimizer. It is a key contribution of this work to provide a general unified convergence analysis for GD optimization methods in the training of DNNs with analytic activations such as the softplus and the popular Gaussian error linear unit (GeLU) activation. Our general unified convergence result applies to a large class of gradient based optimization methods such as the standard GD, the momentum, the Nesterov accelerated gradient (NAG), the RMSprop, the Adam, the Adamax, the Nadam, the Nadamax, the Adan, the AdaBelief, the AMSGrad, and the Yogi optimizers. Our analysis employs the theory of Kurdyka-Łojasiewicz (KL) inequalities to establish convergence to critical points in the training of DNNs. To the best of our knowledge, the generality of our convergence analysis is also just in the special situation of the Adam optimizer a new contribution to the literature on the analysis of AI optimization algorithms.

优化算法深度学习收敛分析Adam

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。