arXiv:2511.11466math.OCcs.LG2025-11被引 8

提出统一理论框架,解释非欧优化为何比传统方法更快。

Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates

  • 基于结构平滑性与梯度噪声假设,统一分析非欧梯度下降。
  • 在稀疏或低秩场景下,收敛速度超越经典欧氏SGD。
  • 可解释SignSGD、Lion等方法优势,适合研究优化算法者阅读。

近期,若干非欧几里得随机梯度下降方法(如SignSGD、Lion、Muon)因其在训练深度神经网络中的实际成功而受到广泛关注。尽管已有诸多工作尝试通过理论分析解释其性能优势,但现有结果无法合理说明这些方法为何优于标准欧氏SGD,因它们的收敛速率并未超越经典欧氏SGD。本文通过建立新的统一收敛性分析,在结构平滑性和梯度噪声假设下解决了这一关键开放问题。结果表明,非欧SGD(i)可利用海森矩阵和梯度噪声上界中的稀疏性或低秩结构;(ii)可严格证明受益于外推法或动量方差缩减等常见技术;(iii)可达到自适应及更复杂算法(如AdaGrad、Shampoo)的最先进收敛速率。

原文摘要 · Abstract (English)

Recently, several instances of non-Euclidean SGD, including SignSGD, Lion, and Muon, have attracted significant interest from the optimization community due to their practical success in training deep neural networks. Consequently, a number of works have attempted to explain this success by developing theoretical convergence analyses. Unfortunately, these results cannot properly justify the superior performance of these methods, as they could not beat the convergence rate of vanilla Euclidean SGD. We resolve this important open problem by developing a new unified convergence analysis under the structured smoothness and gradient noise assumption. In particular, our results indicate that non-Euclidean SGD (i) can exploit the sparsity or low-rank structure of the upper bounds on the Hessian and gradient noise, (ii) can provably benefit from popular algorithmic tools such as extrapolation or momentum variance reduction, and (iii) can match the state-of-the-art convergence rates of adaptive and more complex optimization algorithms such as AdaGrad and Shampoo.

优化算法非欧优化收敛分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。