Muon在凸Lipschitz函数上不收敛,需误差反馈修复但会降效。
Muon Does Not Converge on Convex Lipschitz Functions

- 在凸Lipschitz函数上,无论学习率如何调整,Muon均不收敛。
- 引入误差反馈可恢复收敛性,但导致图像分类与语言建模性能下降。
- 实证成功可能源于光滑性而非凸Lipschitz结构,理论需重构。
Muon及其变体在多种深度学习任务中表现出色。现有对Muon的收敛性分析依赖平滑性假设,但实践中最成功的优化方法(如AdaGrad、Shampoo、Schedule-Free等)主要针对凸且Lipschitz函数设计。本文质疑凸Lipschitz模型是否适用于理解Muon。结果表明:无论学习率调度如何,Muon在凸Lipschitz函数类上均不收敛。此外,误差反馈可恢复含动量的非欧几里得次梯度方法的收敛性,但会使Muon在两个典型场景(CIFAR-10图像分类、nanoGPT on FineWeb-Edu 10B语言建模)中的性能下降。结论是:尽管凸Lipschitz理论在实践方法设计中占据核心地位,却并非理解Muon的合适框架。这意味着Muon的成功可能源于该模型所缺乏的结构,最可能是光滑性。
原文摘要 · Abstract (English)
Muon and its variants have shown strong empirical performance in a variety of deep learning tasks. Existing convergence analyses of Muon rely on smoothness assumptions, though arguably the most successful function class for developing deep learning methods (such as AdaGrad, Shampoo, Schedule-Free and more) has been the class of convex and Lipschitz functions. In this paper we question whether the classical convex Lipschitz model is a useful one for understanding Muon. Our answer is no. We show that Muon does not converge on the class of convex and Lipschitz functions, regardless of the choice of learning rate schedule. We also show that error feedback restores convergence of Muon and all the non-Euclidean subgradient methods with momentum. However, this theoretical fix using error feedback degrades the performance of Muon in two representative settings for image classification (CIFAR-10) and language modeling (nanoGPT on FineWeb-Edu 10B). Our conclusion is that convex Lipschitz theory, despite having a prominent role in the design of practical methods for deep learning, is not the most suited one for Muon. This suggests that Muon's success must come from structure absent from this model, most plausibly related to smoothness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。