arXiv:2510.00643cs.LGmath.OC2025-10被引 12

首个支持压缩通信的非欧优化器,显著减少训练通信量且不损失精度。

Error Feedback for Muon and Friends

  • 基于误差反馈机制,将非欧优化扩展到分布式场景。
  • 实验显示通信量降低7倍,性能与原始方法相当。
  • 适合大规模深度学习分布式训练,尤其关注通信效率的场景。

近期优化器如 Muon、Scion 和 Gluon 通过在非欧范数球上使用逐层线性最小化算子(LMO),突破了大规模深度学习的边界,捕捉到传统算法无法体现的神经网络结构。然而,这类方法尚无严谨的分布式框架,通信瓶颈未被解决。现有少数分布式变体为启发式设计,缺乏收敛保证。本文提出 EF21-Muon,首个具备严格收敛性保证的通信高效、非欧 LMO 优化器。它支持随机梯度、动量及双向压缩与误差反馈,首次将误差反馈拓展至非欧设置。当关闭压缩并选择特定范数时,可还原 Muon/Scion/Gluon。理论覆盖非欧光滑和更一般的 (L^0, L^1)-光滑情形,达到最优欧氏率,并在合适范数下实现更快收敛。进一步分析扩展至逐层(广义)光滑性,捕捉深层网络各向异性结构。在 NanoGPT 基准测试中,相较未压缩的 Muon/Scion/Gluon,EF21-Muon 实现最高达 7× 的通信节省,且无精度下降。

原文摘要 · Abstract (English)

Recent optimizers like Muon, Scion, and Gluon have pushed the frontier of large-scale deep learning by exploiting layer-wise linear minimization oracles (LMOs) over non-Euclidean norm balls, capturing neural network structure in ways traditional algorithms cannot. Yet, no principled distributed framework exists for these methods, and communication bottlenecks remain unaddressed. The very few distributed variants are heuristic, with no convergence guarantees in sight. We introduce EF21-Muon, the first communication-efficient, non-Euclidean LMO-based optimizer with rigorous convergence guarantees. EF21-Muon supports stochastic gradients, momentum, and bidirectional compression with error feedback-marking the first extension of error feedback beyond the Euclidean setting. It recovers Muon/Scion/Gluon when compression is off and specific norms are chosen, providing the first efficient distributed implementation of this powerful family. Our theory covers non-Euclidean smooth and the more general $(L^0, L^1)$-smooth setting, matching best-known Euclidean rates and enabling faster convergence under suitable norm choices. We further extend the analysis to layer-wise (generalized) smoothness regimes, capturing the anisotropic structure of deep networks. Experiments on NanoGPT benchmarking EF21-Muon against uncompressed Muon/Scion/Gluon demonstrate up to $7\times$ communication savings with no accuracy degradation.

非欧优化通信压缩误差反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。