arXiv:2608.12710cs.LGmath.OC2026-08

提出联邦组合优化新方法,提升分布式矩阵模型训练效率与精度。

Federated Compositional Muon Optimizer for Matrix-Wise Models

  • 结合梯度追踪与正交动量,设计新型联邦组合优化器。
  • 理论证明其样本复杂度达O(ε⁻³),优于现有联邦优化算法。
  • 适用于非独立同分布数据的鲁棒联邦学习与风险敏感元学习场景。

Muon 是一种近期发展的优化器,适用于人工智能中的矩阵型模型。尽管已有诸多研究探讨 Muon 及其变体,但这些方法对分层结构问题仍不够适用。为此,我们提出一种高效的联邦组合 Muon(FedCoMuon)优化器,用于求解分布式矩阵型组合优化问题。具体而言,该优化器基于组合梯度追踪与正交化动量机制。此外,我们还提出一种基于动量的方差缩减变体 FedCoMuon-VR。理论上,我们在非独立同分布(non-i.i.d.)和非凸设定下分析了算法的收敛性。特别地,证明 FedCoMuon-VR 在寻找 ε-驻点时的样本复杂度为 O(ε⁻³),低于现有 FedMuon 算法。大量数值实验表明,所提方法在鲁棒联邦学习与任务分布的风险敏感元学习中表现优异,多项指标优于现有组合基线,并在多个设置下达到最佳报告精度。

原文摘要 · Abstract (English)

Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. Although many works have studied Muon and its variants, these methods are still not particularly well-suited for hierarchical structured problems. To fill this gap, we propose an effective federated compositional Muon (FedCoMuon) optimizer to solve distributed matrix-wise compositional optimization problems. Specifically, our FedCoMuon optimizer builds on compositional gradient tracking and orthogonalized momentum. Moreover, we propose a variance reduced variant of FedCoMuon (FedCoMuon-VR) based on a momentum-based variance reduced technique. In theory, we analyze the convergence properties of our algorithms under the non-i.i.d. and non-convex settings. In particular, we prove that our FedCoMuon-VR obtains a lower sample complexity of $O(ε^{-3})$ for finding an $ε$-stationary solution than the existing FedMuon algorithms. Extensive numerical experiments on robust federated learning and task-distributed risk-sensitive meta learning show that our proposed methods are competitive with existing compositional baselines and achieve the best reported accuracy in several settings.

联邦学习优化器组合优化矩阵模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。