arXiv:2511.04456cs.LG2025-11被引 3

解决联邦学习中重尾噪声下的极小极大优化问题,首次给出理论保证。

Federated Stochastic Minimax Optimization under Heavy-Tailed Noises

  • 提出两种新算法:基于归一化梯度和Muon优化器的联合设计。
  • 在较弱条件下实现 $O(1/(TNp)^{(s-1)/(2s)})$ 的收敛速率。
  • 适合存在异常值的分布式非凸极小极大优化场景,如对抗训练。

重尾噪声在非凸随机优化中日益受到关注,因其比标准有界方差假设更贴近实际。本文研究联邦学习中非凸-PL极小极大优化在重尾梯度噪声下的问题。提出两种新算法:Fed-NSGDA-M(融合归一化梯度)与FedMuon-DA(采用Muon优化器进行本地更新)。两者均在较弱条件下有效应对重尾噪声。理论上证明二者均达到 $O(1/(TNp)^{(s-1)/(2s)})$ 的收敛速率。据我们所知,这是首个在重尾噪声下具备严格理论保证的联邦极小极大优化算法。大量实验验证了其有效性。

原文摘要 · Abstract (English)

Heavy-tailed noise has attracted growing attention in nonconvex stochastic optimization, as numerous empirical studies suggest it offers a more realistic assumption than standard bounded variance assumption. In this work, we investigate nonconvex-PL minimax optimization under heavy-tailed gradient noise in federated learning. We propose two novel algorithms: Fed-NSGDA-M, which integrates normalized gradients, and FedMuon-DA, which leverages the Muon optimizer for local updates. Both algorithms are designed to effectively address heavy-tailed noise in federated minimax optimization, under a milder condition. We theoretically establish that both algorithms achieve a convergence rate of $O({1}/{(TNp)^{\frac{s-1}{2s}}})$. To the best of our knowledge, these are the first federated minimax optimization algorithms with rigorous theoretical guarantees under heavy-tailed noise. Extensive experiments further validate their effectiveness.

联邦学习极小极大优化重尾噪声非凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。