研究动量在去中心化联邦学习中的极限,发现其无法克服数据异质性影响。
On the Limits of Momentum in Decentralized and Federated Optimization
- 分析循环参与下的动量机制,揭示其对异质性的敏感性
- 证明步长下降快于1/t时仍收敛至与初始值和异质性相关的常数
- 实验证明该结论适用于真实深度学习场景
近期研究探索了在本地方法中使用动量以提升分布式SGD性能,这在联邦学习(FL)中尤为吸引人,因为动量似乎能缓解统计异质性的影响。然而,尽管已有进展,目前尚不清楚在去中心化场景下(仅部分工作节点每轮参与),动量能否在无界异质性条件下保证收敛。本文分析了循环客户参与下的动量机制,理论上证明其仍不可避免地受统计异质性影响。与SGD类似,我们证明步长衰减策略也无法改善这一问题:任何比Θ(1/t)衰减更快的步长调度,都会导致收敛到一个依赖于初始化和异质性上界的常数值。数值结果验证了理论,深度学习实验也证实了该结论在实际设置中的相关性。
原文摘要 · Abstract (English)
Recent works have explored the use of momentum in local methods to enhance distributed SGD. This is particularly appealing in Federated Learning (FL), where momentum intuitively appears as a solution to mitigate the effects of statistical heterogeneity. Despite recent progress in this direction, it is still unclear if momentum can guarantee convergence under unbounded heterogeneity in decentralized scenarios, where only some workers participate at each round. In this work we analyze momentum under cyclic client participation, and theoretically prove that it remains inevitably affected by statistical heterogeneity. Similarly to SGD, we prove that decreasing step-sizes do not help either: in fact, any schedule decreasing faster than $Θ\left(1/t\right)$ leads to convergence to a constant value that depends on the initialization and the heterogeneity bound. Numerical results corroborate the theory, and deep learning experiments confirm its relevance for realistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。