Muon优化器在重尾噪声下仍能收敛,且比小批量SGD更快。
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
- 通过投影梯度到Stiefel流形保持参数更新正交性。
- 在重尾噪声条件下,证明了其收敛到经验风险的驻点。
- 实验表明其收敛速度优于小批量SGD,适合大规模训练场景。
Muon是一种近期提出的优化器,通过将梯度投影到Stiefel流形以强制参数更新的正交性,在大规模深度神经网络中实现了稳定高效的训练。然而,实际机器学习中的随机噪声常呈现重尾特性,违反了传统方法对有界方差的假设。本文研究了在重尾随机噪声下非凸Hölder光滑经验风险最小化问题,证明在考虑重尾噪声的有界性条件下,Muon可收敛至经验风险的驻点。此外,还表明Muon的收敛速度优于小批量SGD。
原文摘要 · Abstract (English)
Muon is a recently proposed optimizer that enforces orthogonality in parameter updates by projecting gradients onto the Stiefel manifold, leading to stable and efficient training in large-scale deep neural networks. Meanwhile, the previously reported results indicated that stochastic noise in practical machine learning may exhibit heavy-tailed behavior, violating the bounded-variance assumption. In this paper, we consider the problem of minimizing a nonconvex Hölder-smooth empirical risk that works well with the heavy-tailed stochastic noise. We then show that Muon converges to a stationary point of the empirical risk under the boundedness condition accounting for heavy-tailed stochastic noise. In addition, we show that Muon converges faster than mini-batch SGD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。