揭示自适应优化器与非欧几何下降的深层联系,统一理论框架。
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
- 提出自适应光滑性概念,刻画自适应优化器在非凸情形下的收敛机制。
- 证明在特定非欧几何下,自适应优化器可实现加速,而标准光滑性无法做到。
- 引入自适应梯度方差,获得无维度依赖的随机优化收敛保证,适合高维场景。
自适应优化器在仅依据当前梯度调整时,可退化为归一化最速下降(NSD),暗示两者存在紧密关联。关键区别在于分析所依赖的几何结构,例如光滑性定义:凸情形下,自适应优化器受更强的自适应光滑性约束,而NSD使用标准光滑性。本文将自适应光滑性理论拓展至非凸场景,证明其精确刻画了自适应优化器的收敛性。进一步,发现自适应光滑性支持在凸情形下通过Nesterov动量实现加速,而标准光滑性在某些非欧几何下无法提供此类保证。此外,针对随机优化,引入自适应梯度方差,其类比自适应光滑性,并带来不依赖维度的收敛保证,这在标准梯度方差下无法实现。
原文摘要 · Abstract (English)
Adaptive optimizers can reduce to normalized steepest descent (NSD) when only adapting to the current gradient, suggesting a close connection between the two algorithmic families. A key distinction between their analyses, however, lies in the geometries, e.g., smoothness notions, they rely on. In the convex setting, adaptive optimizers are governed by a stronger adaptive smoothness condition, while NSD relies on the standard notion of smoothness. We extend the theory of adaptive smoothness to the nonconvex setting and show that it precisely characterizes the convergence of adaptive optimizers. Moreover, we establish that adaptive smoothness enables acceleration of adaptive optimizers with Nesterov momentum in the convex setting, a guarantee unattainable under standard smoothness for certain non-Euclidean geometry. We further develop an analogous comparison for stochastic optimization by introducing adaptive gradient variance, which parallels adaptive smoothness and leads to dimension-free convergence guarantees that cannot be achieved under standard gradient variance for certain non-Euclidean geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。