统一分析自适应预条件SGD,证明动量可加速其收敛
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
- 构建统一理论框架,覆盖多种自适应优化方法
- 首次证明动量可使AdaGrad类算法超越现有最优收敛率
- 为Adam等算法的高效性提供理论解释,适合优化研究者
本文重新审视带AdaGrad型预条件的随机梯度下降(SGD)。首先,我们建立了在各向异性或矩阵光滑性与噪声假设下的统一收敛性分析,可恢复AdaGrad-Norm、AdaGrad及ASGO/One-sided Shampoo等主流自适应优化方法的最先进收敛结果。此外,揭示了Scion与DASGO两算法间的本质联系,并首次为DASGO提供了理论保证。其次,我们证明通过Nesterov动量,AdaGrad和DASGO等方法的收敛速度可被严格加速,超越现有最佳已知速率。由此首次从理论上证实,AdaGrad类算法可同时受益于对角预条件与动量,或可为Adam的实际效率提供终极解释。
原文摘要 · Abstract (English)
In this paper, we revisit stochastic gradient descent (SGD) with AdaGrad-type preconditioning. Our contributions are twofold. First, we develop a unified convergence analysis of SGD with adaptive preconditioning under anisotropic or matrix smoothness and noise assumptions. This allows us to recover state-of-the-art convergence results for several popular adaptive gradient methods, including AdaGrad-Norm, AdaGrad, and ASGO/One-sided Shampoo. In addition, we establish the fundamental connection between two recently proposed algorithms, Scion and DASGO, and provide the first theoretical guarantees for the latter. Second, we show that the convergence of methods like AdaGrad and DASGO can be provably accelerated beyond the best-known rates using Nesterov momentum. Consequently, we obtain the first theoretical justification that AdaGrad-type algorithms can simultaneously benefit from both diagonal preconditioning and momentum, which may provide an ultimate explanation for the practical efficiency of Adam.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。