arXiv:2602.07425cs.LGcs.CL2026-02被引 12

sign-based优化器在重尾噪声下表现更优,理论首次解释其优势

Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise

  • 基于重尾梯度噪声建模,突破传统方差假设
  • 证明SignSGD和Lion收敛速度优于已有最优界
  • 首次严格分析矩阵优化在重尾随机性下的表现

尽管自适应梯度方法是现代机器学习的核心,但近期SignSGD、Lion和Muon等符号型优化器在训练大语言模型(LLM)时表现出优于AdamW的实证性能。然而,为何符号型更新在方差自适应方法之上仍缺乏理论解释。本文从语言建模中常见的重尾梯度噪声出发,提出一种新的广义重尾噪声条件,比标准有限方差假设更准确刻画LLM行为。在此噪声模型下,我们建立了SignSGD和Lion在广义光滑函数类上的紧收敛速率,达到或超越此前最佳已知界限。进一步,我们对Muon和Muonlight进行了分析,首次提供矩阵优化在重尾随机性下的严格理论支持。实验验证了理论洞察,表明所提噪声模型与实际高度一致。

原文摘要 · Abstract (English)

While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have recently demonstrated superior empirical performance over AdamW in training large language models (LLM). However, a theoretical understanding of why sign-based updates outperform variance-adapted methods remains elusive. In this paper, we aim to bridge the gap between theory and practice through the lens of heavy-tailed gradient noise, a phenomenon frequently observed in language modeling tasks. Theoretically, we introduce a novel generalized heavy-tailed noise condition that captures the behavior of LLMs more accurately than standard finite variance assumptions. Under this noise model, we establish sharp convergence rates of SignSGD and Lion for generalized smooth function classes, matching or surpassing previous best-known bounds. Furthermore, we extend our analysis to Muon and Muonlight, providing what is, to our knowledge, the first rigorous analysis of matrix optimization under heavy-tailed stochasticity. These results offer a strong theoretical justification for the empirical superiority of sign-based optimizers, showcasing that they are naturally suited to handle the noisy gradients associated with heavy tails. Empirically, LLM pretraining experiments validate our theoretical insights and confirm that our proposed noise models are well-aligned with practice.

优化器重尾噪声大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。