Muon优化器可无视维度陷阱,高效逃离鞍点。
Dimension-Free Saddle-Point Escape in Muon

- 通过非线性谱形变机制,打破传统优化器的维数诅咒。
- 理论证明其在足够谱间距下实现定性球状弹射,逃逸时间恒定。
- 适合研究高维非凸优化或设计新型优化算法的研究者。
现代大语言模型训练被极端高维空间中病态平坦的鞍点严重拖慢。针对此问题,我们分析了新兴的Muon优化器在鞍点逃逸中的动态特性,证明其对元素自适应优化器(如AdamW)所受的$/mathcal{O}(D)$维数诅咒具有鲁棒性。通过拓展广义矩阵扰动理论,我们构建了一个捕捉Muon非平衡优化轨迹的理论框架。该框架利用再生核函数演算与宏观Cauchy围道积分,避免了各向同性噪声假设和Tracy-Widom边缘奇异性。我们证明结构非相干性可有效抑制正交漂移,实现维度无关的鞍点逃逸,并在充分谱间隙下触发确定性的$/mathcal{O}(1)$离散弹射。最终,我们给出了Muon的代数维度无关逃逸界,揭示了其非凸优化动力学的本质机制。
原文摘要 · Abstract (English)
Modern Large Language Model (LLM) training is fundamentally bottlenecked by pathologically flat saddle points in extreme high-dimensional landscapes. Motivated by this challenge, we analyze the saddle-point escape dynamics of the emerging Muon optimizer, demonstrating its resilience against the $\mathcal{O}(D)$ dimensional curse that severely traps element-wise adaptive optimizers like AdamW. By extending generalized matrix perturbation theory, we develop a theoretical framework to capture Muon's non-equilibrium optimization trajectories. This theoretical machinery mathematically proves that Muon elegantly bypasses the dimensional curse via a non-linear spectral shaping mechanism. By leveraging resolvent functional calculus and macroscopic Cauchy contour integration, we avoid isotropic noise assumptions and Tracy-Widom edge singularities. We establish that structural incoherence securely shields the trajectory from orthogonal drift, enabling a dimension-free saddle-point escape, and triggering a deterministic $\mathcal{O}(1)$ discrete ballistic ejection under sufficient spectral gap. Consequently, we provide an algebraically dimension-free escape bound for Muon, formalizing the underlying mechanics of its non-convex optimization dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。