arXiv:2505.13416cs.LGmath.OC2025-05被引 64

解决大模型优化器理论与实践脱节问题,提出新算法Gluon。

Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)

  • 基于层间LMO机制设计新优化器,融合理论分析与实际实现。
  • 在纳米GPT和CNN上验证,理论步长与实际调优值高度吻合。
  • 适合关注大模型训练效率与优化器理论的开发者与研究者。

深度学习优化的新进展催生了基于线性最小化预言机(LMO)框架的新型算法,如Muon和Scion。在Adam主导十余年之后,这些方法展现出更好的内存效率、超参数可迁移性以及大规模任务(包括大语言模型训练)上的优异表现,成为有力替代方案。然而,其实际应用与理论理解之间仍存在显著差距:现有分析(1)忽略了实践中层间LMO的应用,(2)依赖不切实际的光滑性假设,导致理论步长过小。为此,本文提出新方法Gluon,可涵盖已有理论方法作为特例,并引入一种新的精细广义光滑性模型,捕捉神经网络的层间几何结构,匹配Muon和Scion的实际分层实现,从而获得具备强实用预测能力的收敛性保证。不同于以往结果,本理论推导出的步长与Pethick等人(2025)报告的微调值高度一致。在NanoGPT和CNN上的实验表明,该假设在优化轨迹中持续成立,最终弥合了理论与实践之间的鸿沟。

原文摘要 · Abstract (English)

Recent developments in deep learning optimization have brought about radically new algorithms based on the Linear Minimization Oracle (LMO) framework, such as $\sf Muon$ and $\sf Scion$. After over a decade of $\sf Adam$'s dominance, these LMO-based methods are emerging as viable replacements, offering several practical advantages such as improved memory efficiency, better hyperparameter transferability, and most importantly, superior empirical performance on large-scale tasks, including LLM training. However, a significant gap remains between their practical use and our current theoretical understanding: prior analyses (1) overlook the layer-wise LMO application of these optimizers in practice, and (2) rely on an unrealistic smoothness assumption, leading to impractically small stepsizes. To address both, we propose a new LMO-based method called $\sf Gluon$, capturing prior theoretically analyzed methods as special cases, and introduce a new refined generalized smoothness model that captures the layer-wise geometry of neural networks, matches the layer-wise practical implementation of $\sf Muon$ and $\sf Scion$, and leads to convergence guarantees with strong practical predictive power. Unlike prior results, our theoretical stepsizes closely match the fine-tuned values reported by Pethick et al. (2025). Our experiments with NanoGPT and CNN confirm that our assumption holds along the optimization trajectory, ultimately closing the gap between theory and practice.

优化器大模型理论分析机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。