证明了局部更新的自适应优化可降低通信开销,提升训练效率。
Convergence of Distributed Adaptive Optimization with Local Updates
- 提出新分析方法,证明局部SGDM和Local Adam在特定条件下优于传统方法。
- 在凸与弱凸场景下,局部更新能减少通信次数,且收敛性更优。
- 适用于大规模分布式训练,尤其适合通信成本高的场景。
我们研究了具有局部更新(间歇通信)的分布式自适应算法。尽管自适应方法在现代机器学习模型的分布式训练中表现出色,但其局部更新在降低通信复杂度方面的理论优势尚未完全明确。本文首次证明,在某些情况下,带动量的局部随机梯度下降(Local SGDM)和局部Adam(Local Adam)在凸与弱凸设置下可优于其批量版本。我们的分析基于一种新技巧,用于证明局部迭代过程中的收缩性,这是展现局部更新优势的关键步骤,该分析在广义光滑性假设和梯度裁剪策略下成立。
原文摘要 · Abstract (English)
We study distributed adaptive algorithms with local updates (intermittent communication). Despite the great empirical success of adaptive methods in distributed training of modern machine learning models, the theoretical benefits of local updates within adaptive methods, particularly in terms of reducing communication complexity, have not been fully understood yet. In this paper, for the first time, we prove that \em Local SGD \em with momentum (\em Local \em SGDM) and \em Local \em Adam can outperform their minibatch counterparts in convex and weakly convex settings in certain regimes, respectively. Our analysis relies on a novel technique to prove contraction during local iterations, which is a crucial yet challenging step to show the advantages of local updates, under generalized smoothness assumption and gradient clipping strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。