arXiv:2603.18899cs.LGmath.OC2026-03被引 4

首次为Adam优化器提供无条件误差分析,解决其理论缺陷

Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method

  • 建立Adam的统一先验界,突破以往依赖收敛假设的局限
  • 证明在强凸随机优化问题中,Adam误差可被严格控制
  • 适合研究优化算法理论或深度学习训练稳定性的学者

自适应矩估计(Adam)优化器是人工智能系统中训练深度神经网络最流行的随机梯度下降方法之一。尽管在实际应用中取得巨大成功,但对其完整的误差分析仍是开放性问题,尤其在处理强凸随机优化问题(SOP)时。已有研究的误差分析多基于假设:Adam不会发散至无穷大,即保持一致有界。本文的关键贡献在于首次为Adam建立了统一的先验界,从而实现了对一大类强凸随机优化问题的无条件误差分析。

原文摘要 · Abstract (English)

The adaptive moment estimation (Adam) optimizer proposed by Kingma & Ba (2014) is presumably the most popular stochastic gradient descent (SGD) optimization method for the training of deep neural networks (DNNs) in artificial intelligence (AI) systems. Despite its groundbreaking success in the training of AI systems, it still remains an open research problem to provide a complete error analysis of Adam, not only for optimizing DNNs but even when applied to strongly convex stochastic optimization problems (SOPs). Previous error analysis results for strongly convex SOPs in the literature provide conditional convergence analyses that rely on the assumption that Adam does not diverge to infinity but remains uniformly bounded. It is the key contribution of this work to establish uniform a priori bounds for Adam and, thereby, to provide -- for the first time -- an unconditional error analysis for Adam for a large class of strongly convex SOPs.

优化算法误差分析Adam

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。