arXiv:2505.23737stat.MLcs.IT2025-05

Muon优化器利用矩阵结构提升训练效率,理论证明其在低秩海森结构下优于传统方法。

On the Convergence Analysis of Muon

  • 针对矩阵参数设计专用优化器Muon,保留其固有结构特性
  • 理论证明在低秩海森结构下,Muon收敛速度优于梯度下降
  • 适用于具有低秩特征的神经网络训练,如大模型微调

神经网络中多数参数天然以矩阵形式存在,但主流优化器将其视为扁平向量处理,可能忽略其结构特性。近期提出的Muon优化器专为矩阵参数设计,实验表明其在训练神经网络时显著优于传统方法。然而,其收敛行为与性能优势的理论基础仍不明确。本文对Muon进行了全面的收敛速率分析,并与梯度下降(GD)对比,揭示了其在低秩海森矩阵条件下可实现更优性能的理论依据。该现象在实际神经网络训练中广泛存在。实验结果验证并支持了理论发现。

原文摘要 · Abstract (English)

The majority of parameters in neural networks are naturally represented as matrices. However, most commonly used optimizers treat these matrix parameters as flattened vectors during optimization, potentially overlooking their inherent structural properties. Recently, an optimizer called Muon has been proposed, specifically designed to optimize matrix-structured parameters. Extensive empirical evidence shows that Muon can significantly outperform traditional optimizers when training neural networks. Nonetheless, the theoretical understanding of Muon's convergence behavior and the reasons behind its superior performance remain limited. In this work, we present a comprehensive convergence rate analysis of Muon and its comparison with Gradient Descent (GD). We characterize the conditions under which Muon can outperform GD. Our theoretical results reveal that Muon can benefit from the low-rank structure of Hessian matrices, a phenomenon widely observed in practical neural network training. Our experimental results support and corroborate the theoretical findings.

优化器矩阵结构收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。