改进了缪子优化器在非凸优化中的收敛速度理论分析
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
- 采用简化直接分析法,无需限制性假设
- 收敛速度更快,适用问题范围更广
- 为正交化一阶优化方法提供理论参考
缪子优化器因其正交化的一阶更新机制而受到关注,对其收敛行为的深入理论理解对指导实际应用至关重要;然而,现有收敛保证要么过于粗略,要么依赖于严格的分析假设。本文通过一种不依赖更新规则限制性假设的直接且简化的分析,建立了缪子优化器更精确的收敛保证。结果在保持更广泛问题设置覆盖的同时,实现了更快的收敛速率,显著优于现有边界。这些发现提供了对缪子优化器更准确的理论刻画,并为一类更广泛的正交化一阶优化方法提供了可借鉴的洞察。
原文摘要 · Abstract (English)
The Muon optimizer has recently attracted attention due to its orthogonalized first-order updates, and a deeper theoretical understanding of its convergence behavior is essential for guiding practical applications; however, existing convergence guarantees are either coarse or obtained under restrictive analytical settings. In this work, we establish sharper convergence guarantees for the Muon optimizer through a direct and simplified analysis that does not rely on restrictive assumptions on the update rule. Our results improve upon existing bounds by achieving faster convergence rates while covering a broader class of problem settings. These findings provide a more accurate theoretical characterization of Muon and offer insights applicable to a broader class of orthogonalized first-order methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。