在简单矩阵分解任务中,Muon并不总优于AdamW,挑战其在大模型中的优势归因。
Reassessing Muon for Matrix Factorization

- 用低秩矩阵分解控制变量,剥离大模型复杂性影响
- 对比发现Muon在该任务中未持续胜过调优后的AdamW
- 提醒需在可控问题上评估优化器,而非仅依赖端到端表现
Muon最近被提出作为大规模深度学习的强效优化器,通过近似正交化重塑梯度更新,在大语言模型训练中报告优于Adam和AdamW。其成功引发理论研究,将其解释为谱范数下的最速下降。然而,目前尚不清楚其优势究竟来自更新规则本身,还是现代深度网络的规模、架构与数据的产物。本文通过在简单、可理解且谱结构明确的问题——低秩矩阵因子分解上研究Muon,将优化器与这些混淆因素分离。通过与精心调参的自适应基线进行对照,发现Muon在此设置下并未持续优于AdamW,且多个先前报告的优势对超参数选择敏感。结果揭示了谱感知正交化在何种条件下有益,并主张在端到端基准之外,也应在受控问题上评估现代优化器。
原文摘要 · Abstract (English)
Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm. Yet it remains unclear which of Muon's advantages stem from its update rule itself and which are artifacts of the scale, architecture, and data of modern deep networks. In this work, we isolate the optimizer from these confounding factors by studying Muon on a simple, well-understood, and spectrally structured problem: low-rank matrix factorization. Through a controlled comparison against carefully tuned adaptive baselines, we find that Muon does not consistently outperform AdamW in this setting and that several previously reported advantages are sensitive to hyperparameter choices. Our results provide a more nuanced picture of when spectrum-aware orthogonalization is beneficial and argue for evaluating modern optimizers on controlled problems in addition to end-to-end benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。