arXiv:2505.04005cs.LG2025-05

发现大型模型中muon优化器的正交化过程会因矩阵奇异值缩小而失效。

Iterative Orthogonalization Scaling Laws

  • 分析muon优化器在大规模下的迭代正交化机制
  • 理论与实验验证随机矩阵奇异值随规模缩小
  • 揭示大模型训练中的潜在数值问题,适合研究优化器者参考

近期,muon优化器因其作为广泛使用的Adam优化器替代品的潜力而受到关注。尽管已有研究记录了muon在不同超参数下的缩放规律,如权重衰减和学习率,但在更大规模下,muon中包含的迭代正交化过程可能面临一个问题:随着规模增大,随机矩阵的奇异值会缩小。本文从理论上和实验上验证了这一缩放行为,但未提出解决方法。

原文摘要 · Abstract (English)

The muon optimizer has picked up much attention as of late as a possible replacement to the seemingly omnipresent Adam optimizer. Recently, care has been taken to document the scaling laws of hyper-parameters under muon such as weight decay and learning rate. However, at much larger scales the iterative orthogonalization procedure present in muon may suffer a possible issue as the singular values of random matrices shrink with scale. This paper shows this scaling behavior theoretically and empirically on random matrices but does not suggest what to do about it.

优化器缩放定律数值稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。