arXiv:2509.11254math.OCcs.LG2025-09

提出改进版PowerSGD+,解决梯度压缩不收敛问题

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees

  • 通过定期更新投影子空间,保持与最优方向对齐
  • 证明了新方法在标准假设下可保证收敛
  • 适合大规模语言模型分布式训练场景

低秩梯度压缩方法如PowerSGD在通信高效分布式优化中受到关注。然而,PowerSGD在随机设置下的收敛性仍不明确。本文指出PowerSGD并不总能收敛到最优解,并提供明确反例支持该结论。为此,我们提出PowerSGD+,通过奇异值分解定期更新投影子空间,确保其始终与最优子空间对齐。我们证明了PowerSGD+在标准假设下可收敛,并通过大规模语言模型任务的实证评估验证了其有效性。

原文摘要 · Abstract (English)

Low-rank gradient compression methods, such as PowerSGD, have gained attention in communication-efficient distributed optimization. However, the convergence guarantees of PowerSGD remain unclear, particularly in stochastic settings. In this paper, we show that PowerSGD does not always converge to the optimal solution and provide a clear counterexample to support this finding. To address this, we introduce PowerSGD+, which periodically updates the projection subspace via singular value decomposition, ensuring that it remains aligned with the optimal subspace. We prove that PowerSGD+ converges under standard assumptions and validate its effectiveness through empirical evaluation on large language model tasks.

分布式训练梯度压缩收敛性保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。