arXiv:2505.02222cs.LGstat.ML2025-05被引 69

Muon优化器在大批次下比AdamW更高效,兼顾训练效果与计算成本。

Practical Efficiency of Muon for Pretraining

  • 使用二阶优化器的简化版本Muon,提升大批次训练效率。
  • 在超临界批量下仍保持数据效率,40亿参数模型验证有效。
  • 兼容muP参数化,支持超参迁移,资源开销小。

我们证明,Muon——一种最简化的二阶优化器,明确扩展了AdamW在计算-时间权衡上的帕累托前沿。实验发现,当批量远超所谓临界批量时,Muon在保持数据效率方面优于AdamW,同时维持计算高效性,从而实现更经济的训练。我们研究了Muon与最大更新参数化(muP)的结合,用于高效超参迁移,并提出一种简单望远镜算法,可解释muP中所有误差来源,仅引入轻微资源开销。通过最大达四亿参数模型的广泛实验,以及对数据分布和架构的消融分析,验证了结论的有效性。

原文摘要 · Abstract (English)

We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon is more effective than AdamW in retaining data efficiency at large batch sizes, far beyond the so-called critical batch size, while remaining computationally efficient, thus enabling more economical training. We study the combination of Muon and the maximal update parameterization (muP) for efficient hyperparameter transfer and present a simple telescoping algorithm that accounts for all sources of error in muP while introducing only a modest overhead in resources. We validate our findings through extensive experiments with model sizes up to four billion parameters and ablations on the data distribution and architecture.

优化器大模型训练高效训练二阶优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。