arXiv:2603.15059cs.LGmath.OC2026-03被引 1

Muon优化器在重尾噪声下仍能收敛,且比小批量SGD更快。

Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization

  • 通过投影梯度到Stiefel流形保持参数更新正交性。
  • 在重尾噪声条件下,证明了其收敛到经验风险的驻点。
  • 实验表明其收敛速度优于小批量SGD,适合大规模训练场景。

Muon是一种近期提出的优化器,通过将梯度投影到Stiefel流形以强制参数更新的正交性,在大规模深度神经网络中实现了稳定高效的训练。然而,实际机器学习中的随机噪声常呈现重尾特性,违反了传统方法对有界方差的假设。本文研究了在重尾随机噪声下非凸Hölder光滑经验风险最小化问题,证明在考虑重尾噪声的有界性条件下,Muon可收敛至经验风险的驻点。此外,还表明Muon的收敛速度优于小批量SGD。

原文摘要 · Abstract (English)

Muon is a recently proposed optimizer that enforces orthogonality in parameter updates by projecting gradients onto the Stiefel manifold, leading to stable and efficient training in large-scale deep neural networks. Meanwhile, the previously reported results indicated that stochastic noise in practical machine learning may exhibit heavy-tailed behavior, violating the bounded-variance assumption. In this paper, we consider the problem of minimizing a nonconvex Hölder-smooth empirical risk that works well with the heavy-tailed stochastic noise. We then show that Muon converges to a stationary point of the empirical risk under the boundedness condition accounting for heavy-tailed stochastic noise. In addition, we show that Muon converges faster than mini-batch SGD.

优化器非凸优化重尾噪声收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。