arXiv:2503.08427math.OCcs.LG2025-03被引 4

提出ADEF算法,让分布式训练更快更省通信。

Accelerated Distributed Optimization with Compression and Error Feedback

  • 结合加速、压缩与误差反馈,降低通信开销。
  • 理论证明在凸优化下达到首个加速收敛率。
  • 适合大规模分布式训练场景,尤其关注通信效率的工程师。

现代机器学习任务通常涉及海量数据和模型,需要通信开销更低的分布式优化算法。通信压缩(客户端向服务器传输压缩更新)已成为缓解通信瓶颈的关键技术。然而,对于带有压缩的随机分布式优化,特别是与Nesterov加速结合时,理论理解仍不充分。本文提出新算法ADEF(加速分布式误差反馈),融合Nesterov加速、压缩、误差反馈及梯度差压缩。我们证明ADEF在一般凸情形下首次实现了带压缩的随机分布式优化的加速收敛率。数值实验验证了理论结果,并展示了ADEF在降低通信成本的同时保持快速收敛的实际有效性。

原文摘要 · Abstract (English)

Modern machine learning tasks often involve massive datasets and models, necessitating distributed optimization algorithms with reduced communication overhead. Communication compression, where clients transmit compressed updates to a central server, has emerged as a key technique to mitigate communication bottlenecks. However, the theoretical understanding of stochastic distributed optimization with contractive compression remains limited, particularly in conjunction with Nesterov acceleration -- a cornerstone for achieving faster convergence in optimization. In this paper, we propose a novel algorithm, ADEF (Accelerated Distributed Error Feedback), which integrates Nesterov acceleration, contractive compression, error feedback, and gradient difference compression. We prove that ADEF achieves the first accelerated convergence rate for stochastic distributed optimization with contractive compression in the general convex regime. Numerical experiments validate our theoretical findings and demonstrate the practical efficacy of ADEF in reducing communication costs while maintaining fast convergence.

分布式优化压缩通信加速算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。