arXiv:2410.04458cs.LGmath.OC2024-10ICML被引 12

提出新框架证明Adam在更宽松条件下收敛,逼近SGD理论表现。

A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD

  • 构建新分析框架,放宽对梯度的约束假设
  • 在SGD常用条件下实现渐近收敛与相似样本复杂度
  • 为Adam提供更贴近实际应用的理论支撑,适合研究者参考

自适应矩估计(Adam)是深度学习中的核心优化算法,因其自适应学习率和高效处理大规模数据的能力而广受认可。然而,其收敛性的理论理解长期受限于严格假设,如几乎必然有界的随机梯度或统一有界梯度,这些条件比分析随机梯度下降(SGD)所需的更为苛刻。本文提出一个新颖且全面的分析框架,用于研究Adam的收敛性。该框架表明,在通常用于分析SGD的较弱假设下——即L-光滑性和ABC不等式——Adam在几乎必然意义和L₁意义下均能实现渐近收敛(最后一迭代意义)。同时,在相同假设下,Adam可达到与SGD相当的非渐近样本复杂度界。

原文摘要 · Abstract (English)

Adaptive Moment Estimation (Adam) is a cornerstone optimization algorithm in deep learning, widely recognized for its flexibility with adaptive learning rates and efficiency in handling large-scale data. However, despite its practical success, the theoretical understanding of Adam's convergence has been constrained by stringent assumptions, such as almost surely bounded stochastic gradients or uniformly bounded gradients, which are more restrictive than those typically required for analyzing stochastic gradient descent (SGD). In this paper, we introduce a novel and comprehensive framework for analyzing the convergence properties of Adam. This framework offers a versatile approach to establishing Adam's convergence. Specifically, we prove that Adam achieves asymptotic (last iterate sense) convergence in both the almost sure sense and the \(L_1\) sense under the relaxed assumptions typically used for SGD, namely \(L\)-smoothness and the ABC inequality. Meanwhile, under the same assumptions, we show that Adam attains non-asymptotic sample complexity bounds similar to those of SGD.

优化算法收敛分析Adam

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。