揭示去中心化训练中网络拓扑对收敛速度的真实影响机制
Improved Convergence Analysis of Topology Dependence in Decentralized SGD
- 突破传统仅用谱间隙分析,全面考虑混合矩阵所有特征值的影响
- 实验证明拓扑选择在异构场景下显著影响收敛速度
- 为去中心化学习的拓扑设计提供更精准的理论指导
去中心化随机梯度下降(Decentralized SGD)是分布式学习中的基础算法,但其收敛行为受底层网络拓扑的影响尚未被充分理解。现有分析表明,谱间隙较小的拓扑会显著降低同质与异质情况下的收敛速率。然而,以往实验显示拓扑选择在异质场景中影响显著,而在同质场景中影响甚微。本文提出更紧致的收敛分析,揭示混合矩阵的所有特征值均影响收敛速度,而不仅限于谱间隙。通过系统实验验证,新分析能更准确描述拓扑对收敛速率的影响,为去中心化学习的拓扑设计提供更精细的理论依据。
原文摘要 · Abstract (English)
Decentralized SGD is a fundamental algorithm in decentralized learning, although the influence of an underlying network topology on its convergence behavior is not yet fully understood. Existing convergence analyses have shown that topologies with a small spectral gap significantly deteriorate the convergence rate of Decentralized SGD in both homogeneous and heterogeneous cases. However, many prior papers have reported that indeed the choice of the topology has a significant experimental impact in the heterogeneous case, but has little experimental impact on training behavior in the homogeneous case. In this paper, we present a tighter convergence analysis of Decentralized SGD, offering a more precise understanding of how topologies affect the convergence rate than the prior analysis. Specifically, unlike existing convergence analyses that used only the spectral gap as a property of the topology, our novel analysis shows that all eigenvalues of the mixing matrix affect the convergence rate. Throughout the experiments, we carefully evaluated the convergence behavior of Decentralized SGD and demonstrated that our novel convergence analysis can more accurately describe the effect of topology on the convergence rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。