arXiv:2505.10322cs.LGmath.OC2025-05被引 2

提出抗延迟的异步分布式优化算法,提升实际场景下的学习效率。

Asynchronous Decentralized SGD under Non-Convexity: A Block-Coordinate Descent Framework

  • 基于块坐标下降框架分析异步算法收敛性
  • 无需假设数据异质性,可在任意延迟下收敛
  • 适合计算/通信不均衡的真实分布式系统

去中心化优化在无中心控制下利用分布式数据方面日益重要,可提升可扩展性和隐私性。然而实际部署面临计算速度异质和通信延迟不可预测等挑战。本文在计算与通信时间有界的合理假设下,提出了异步去中心化随机梯度下降(ADSGD)的改进模型。为理解其收敛性,先分析了异步随机块坐标下降(ASBCD)作为工具,随后证明ADSGBD在计算延迟无关的步长下仍能收敛,且无需假设数据异质性有界。实验表明,ADSGBD在多种场景下均优于现有方法,在墙钟时间上的收敛速度更快。该方法具有内存与通信开销小、对通信和计算延迟鲁棒等优势,适用于真实世界的去中心化学习任务。

原文摘要 · Abstract (English)

Decentralized optimization has become vital for leveraging distributed data without central control, enhancing scalability and privacy. However, practical deployments face fundamental challenges due to heterogeneous computation speeds and unpredictable communication delays. This paper introduces a refined model of Asynchronous Decentralized Stochastic Gradient Descent (ADSGD) under practical assumptions of bounded computation and communication times. To understand the convergence of ADSGD, we first analyze Asynchronous Stochastic Block Coordinate Descent (ASBCD) as a tool, and then show that ADSGD converges under computation-delay-independent step sizes. The convergence result is established without assuming bounded data heterogeneity. Empirical experiments reveal that ADSGD outperforms existing methods in wall-clock convergence time across various scenarios. With its simplicity, efficiency in memory and communication, and resilience to communication and computation delays, ADSGD is well-suited for real-world decentralized learning tasks.

去中心化学习异步优化分布式训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。