提出适用于非凸联邦学习的新算法,无需额外假设。
Methods with Local Steps and Random Reshuffling for Generally Smooth Non-Convex Federated Optimization
- 结合本地步长与随机重排,优化客户端与服务器步长比例。
- 在广义光滑性下实现理论收敛,实验验证有效。
- 适合无额外约束的非凸联邦学习场景,研究者必读。
非凸机器学习问题通常不满足标准光滑性假设。基于经验发现,Zhang 等(2020b)提出了更现实的广义 (L₀, L₁)-光滑性假设,但该领域仍基本未被探索。现有针对标准光滑性设计的算法需重新调整。然而,在联邦学习背景下,仅少数工作关注此问题,且依赖额外限制性假设。本文填补这一空白:提出并分析了新方法,支持本地步骤、部分客户端参与及随机重排,且仅依赖广义光滑性,无额外限制假设。所提方法通过客户端与服务器步长的合理配合及梯度裁剪实现。此外,首次在 Polyak-Łojasiewicz 条件下分析了这些方法。理论结果与标准光滑性下的已知结论一致,实验结果支持理论洞察。
原文摘要 · Abstract (English)
Non-convex Machine Learning problems typically do not adhere to the standard smoothness assumption. Based on empirical findings, Zhang et al. (2020b) proposed a more realistic generalized $(L_0, L_1)$-smoothness assumption, though it remains largely unexplored. Many existing algorithms designed for standard smooth problems need to be revised. However, in the context of Federated Learning, only a few works address this problem but rely on additional limiting assumptions. In this paper, we address this gap in the literature: we propose and analyze new methods with local steps, partial participation of clients, and Random Reshuffling without extra restrictive assumptions beyond generalized smoothness. The proposed methods are based on the proper interplay between clients' and server's stepsizes and gradient clipping. Furthermore, we perform the first analysis of these methods under the Polyak-Ł ojasiewicz condition. Our theory is consistent with the known results for standard smooth problems, and our experimental results support the theoretical insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。