arXiv:2510.03866cs.LGstat.ML2025-10被引 8

提出FedMuon优化器,提升联邦学习收敛性与鲁棒性。

On Provable Benefits of Muon in Federated Learning

  • 设计新算法FedMuon,利用正交化更新方向改进联邦学习
  • 理论证明非凸问题下收敛速度,且学习率与参数无关
  • 天然适配重尾噪声,适合异构数据场景下的联邦学习

最近提出的优化器Muon在多种应用中表现出色,但在联邦学习中的有效性尚未研究。本文针对此空白,提出新算法FedMuon,并建立了其在非凸问题下的收敛速率。理论分析表明,由于更新方向正交化,FedMuon的学习率独立于具体问题参数,且能自然处理重尾噪声。大量实验在多种神经网络架构上验证了该算法的有效性。

原文摘要 · Abstract (English)

The recently introduced optimizer, Muon, has gained increasing attention due to its superior performance across a wide range of applications. However, its effectiveness in federated learning remains unexplored. To address this gap, this paper investigates the performance of Muon in the federated learning setting. Specifically, we propose a new algorithm, FedMuon, and establish its convergence rate for nonconvex problems. Our theoretical analysis reveals multiple favorable properties of FedMuon. In particular, due to its orthonormalized update direction, the learning rate of FedMuon is independent of problem-specific parameters, and, importantly, it can naturally accommodate heavy-tailed noise. The extensive experiments on a variety of neural network architectures validate the effectiveness of the proposed algorithm.

联邦学习优化器正交化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。