Muon优化器的性能突破,源于对梯度极分解的近似处理带来的动态优势。
Insights on Muon from Simple Quadratics
- 通过分析简单二次函数,揭示了极分解近似对离散动态的质变影响
- 近似误差反而提升收敛可达性与有限步数性能,非单纯精度损失
- 目标函数结构影响优化常数,现有条件数解释不足
Muon 在大规模训练中表现优异,其更新方式基于梯度的(近似)极分解。现有理论多聚焦于单步比较(基于二次代理)和最坏情况保证,将极分解的不精确性视为需被“排除”的干扰因素。本文表明,在如 $L(W)=\frac12\|W\|_{\text{F}}^2$ 这类简单强凸函数上,这些视角已显不足,说明理解 Muon 需超越局部代理与悲观最坏情况分析。我们的分析揭示两个关键现象:(i) 极分解步骤中的近似误差可定性改变离散时间动力学,提升可达性与有限时间性能——这正是实践者调参所依赖的效果,但现有理论多将其视为纯粹精度妥协;(ii) 目标函数的结构性质会影响有限预算下的常数,超出传统条件数解释范围。因此,任何涵盖此类情况的通用理论必须显式包含这些要素,或说明它们在感兴趣区间内无关。
原文摘要 · Abstract (English)
Muon updates weight matrices along (approximate) polar factors of the gradients and has shown strong empirical performance in large-scale training. Existing attempts at explaining its performance largely focus on single-step comparisons (on quadratic proxies) and worst-case guarantees that treat the inexactness of the polar-factor as a nuisance ``to be argued away''. We show that already on simple strongly convex functions such as $L(W)=\frac12\|W\|_{\text{F}}^2$, these perspectives are insufficient, suggesting that understanding Muon requires going beyond local proxies and pessimistic worst-case bounds. Instead, our analysis exposes two observations that already affect behavior on simple quadratics and are not well captured by prevailing abstractions: (i) approximation error in the polar step can qualitatively alter discrete-time dynamics and improve reachability and finite-time performance -- an effect practitioners exploit to tune Muon, but that existing theory largely treats as a pure accuracy compromise; and (ii) structural properties of the objective affect finite-budget constants beyond the prevailing conditioning-based explanations. Thus, any general theory covering these cases must either incorporate these ingredients explicitly or explain why they are irrelevant in the regimes of interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。