解析了非光滑优化中谱下降法的收敛性,为大模型训练提供理论支持。
Convergence of Spectral Descent for Non-smooth Optimization
- 提出谱下降与截断谱下降,适用于非光滑凸优化问题。
- 在凸、利普希茨连续和尖锐条件下,证明全局线性收敛。
- 适用于低秩矩阵恢复等实际任务,适合研究优化算法的学者。
Muon优化器在训练大语言模型方面展现出显著的实证效果,但其理论机制仍不清晰。现有对Muon的收敛性保证依赖于光滑性假设,其在非光滑情况下的行为尚无系统研究。本文通过分析Muon的简化版本——谱下降(SD)及其截断形式(TSD),在凸性、利普希茨连续性和尖锐性条件下,建立了二者在非光滑凸优化中的全局线性收敛性。此外,研究了带有解耦权重衰减的正则化变体,并通过与Frank-Wolfe方法的关联,获得次线性收敛保证。最后,将理论框架应用于混合稀疏与密集噪声下的鲁棒低秩矩阵恢复,给出了严格的恢复保证。数值实验验证了理论结果,并展示了Muon类方法在非光滑优化中的有效性。
原文摘要 · Abstract (English)
The Muon optimizer has recently demonstrated remarkable empirical success in training large language models. However, the theoretical understanding of its mechanisms remains limited. Current convergence guarantees for Muon rely heavily on smoothness assumptions, leaving its non-smooth convergence behavior largely unexplored. In this work, we take a step toward bridging this gap by investigating Spectral Descent (SD), a simplified variant of Muon, together with its truncated counterpart, Truncated Spectral Descent (TSD). Under convexity, Lipschitz continuity, and sharpness conditions, we establish global linear convergence for both SD and TSD in non-smooth convex formulations. We also study regularized variants equipped with decoupled weight decay and derive sublinear convergence guarantees through their connection with Frank-Wolfe methods. Finally, we apply our theoretical framework to robust low-rank matrix recovery under mixed sparse and dense noise regimes and provide rigorous recovery guarantees. Numerical experiments support the theoretical findings and demonstrate the effectiveness of Muon-type methods for non-smooth optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。