arXiv:2506.00961cs.LGstat.ML2025-06ICML被引 1

提出新算法,让分布式学习能用更多机器且不降速

Enhancing Parallelism in Decentralized Stochastic Convex Optimization

  • 设计可随时运行的分布式优化算法,突破并行上限
  • 理论证明能支持更大规模网络,收敛速度优于现有方法
  • 适合大规模分布式训练场景,尤其是高连接拓扑

去中心化学习已成为跨多台机器高效处理大数据的有力方法。然而,这类方法常受限于扩展性:当机器数超过一定阈值后,收敛速度会下降。本文提出一种新型去中心化学习算法——去中心化任意时间SGD,显著提升了关键并行度阈值,使更多机器可被有效利用而性能不降。在随机凸优化框架下,我们建立了超越当前最优的并行性理论上限,使更大规模网络在高度连通拓扑中仍能获得良好统计保证,缩小了与集中式学习的差距。

原文摘要 · Abstract (English)

Decentralized learning has emerged as a powerful approach for handling large datasets across multiple machines in a communication-efficient manner. However, such methods often face scalability limitations, as increasing the number of machines beyond a certain point negatively impacts convergence rates. In this work, we propose Decentralized Anytime SGD, a novel decentralized learning algorithm that significantly extends the critical parallelism threshold, enabling the effective use of more machines without compromising performance. Within the stochastic convex optimization (SCO) framework, we establish a theoretical upper bound on parallelism that surpasses the current state-of-the-art, allowing larger networks to achieve favorable statistical guarantees and closing the gap with centralized learning in highly connected topologies.

分布式优化去中心化随机凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。