arXiv:2606.23676cs.LGcs.AI2026-06被引 1

探究AdamW在重尾噪声下的有效性,揭示其收敛机制的理论空白

Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

  • 提出重尾噪声下AdamW收敛性的开放问题
  • 证明加权度量基准为正,支持其潜在有效性
  • 揭示二阶矩累积器可能隐藏大梯度的下界机制

AdamW是训练大语言模型(LLMs)的默认优化器,但其理论基础仍局限于有限方差场景。这愈发令人不满,因为实证表明大模型预训练中的随机梯度噪声通常具有重尾特性。近期研究表明,基于符号的优化器(如Lion、Muon)可实现尖锐的重尾收敛速率,且AdaGrad也能在重尾噪声下收敛。然而,目前尚无关于AdamW在该设定下的严格收敛理论。本文将此问题形式化为一个开放问题,证明了一个正向的加权度量基准,并提出一种围道下界机制,说明分母记忆可能隐藏大梯度。

原文摘要 · Abstract (English)

AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This is increasingly unsatisfying, as empirical evidence indicates that stochastic gradient noise in LLM pretraining is typically heavy-tailed. Recent work shows that sign-based optimizers such as Lion and Muon achieve sharp heavy-tailed rates, and that AdaGrad can also converge under heavy-tailed noise. However, no rigorous convergence theory for AdamW has yet been established in this regime. Can AdamW converge under the same heavy-tailed assumptions, or does its second-moment accumulator create a genuine obstruction? We formulate this as an open problem, prove a positive weighted-metric benchmark, and give a corridor lower-bound mechanism showing how denominator memory can hide large gradients.

优化器重尾噪声收敛性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。