arXiv:2606.09658cs.LGcs.AI2026-06被引 1

Muon比Adam更鲁棒且迁移性更强,适合预训练模型

Muon Learns More Robust and Transferable Features than Adam

论文配图:Muon Learns More Robust and Transferable Features than Adam
图 1 · 摘自论文原文
  • 用层间探针分析发现Muon学习到的特征具有更大逻辑值间隔
  • 在下游任务上,Muon特征迁移效果优于Adam和SGD
  • 理论证明其在多成分特征中具有更大间隔和更高有效秩

Muon作为前沿优化器,在预训练大语言模型和视觉分类器中表现优异。本文从鲁棒性和迁移性角度研究其特征学习优势。通过在带噪声图像和文本上评估预训练模型,发现Muon在不同架构(包括Transformer和CNN)下均比Adam和SGD更具鲁棒性。利用层间探针分析显示,该优势体现在各层更大的逻辑值间隔。进一步在下游任务中训练线性分类器或微调全模型,验证了Muon特征具有更强迁移能力,其隐藏状态多样性(以有效秩衡量)也更高。在含多成分特征的典型分类问题中,理论上证明Muon获得更大间隔和更高有效秩,支持了实证结果。

原文摘要 · Abstract (English)

Muon has recently emerged as a state-of-the-art optimizer for pretraining Large Language Models (LLMs) and vision classifiers. Despite its efficiency advantage over Adam and SGD, the feature-learning advantage of Muon remains unclear. This paper investigates Muon's feature-learning advantage through the lens of robustness and transferability. First, by evaluating pretrained models on corrupted images and texts, we show that features learned by Muon are consistently more robust than those learned by Adam and SGD across different architectures, including transformers and Convolutional Neural Networks (CNNs). Using trained layer-wise probes, we further show that this robustness advantage is reflected in larger logit margins across layers. Second, by training linear classifiers or fine-tuning full models from pretrained parameters on downstream tasks, we demonstrate that Muon-learned features transfer more effectively than those learned by Adam and SGD. This transferability advantage is further supported by the diversity of hidden states across layers, as measured by effective rank. Finally, in a representative classification problem with multi-component features, we prove that Muon attains larger margins and higher effective rank than Adam and SGD, providing theoretical support for our empirical findings.

优化器特征鲁棒性迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。