arXiv:2510.22539eess.SYcs.LG2025-10NeurIPS被引 1

针对异构慢节点问题,提出优化编码方案提升分布式学习效率。

Approximate Gradient Coding for Distributed Learning with Heterogeneous Stragglers

  • 基于个体慢节点概率设计最优编码与解码系数
  • 显著降低误差并加速收敛,优于现有方法
  • 适合大规模异构分布式训练场景

本文提出一种优化结构的梯度编码方案,以缓解分布式学习中的慢节点问题。传统梯度编码方法常假设同质慢节点模型或依赖过度数据冗余,在真实异构系统中性能受限。为此,我们建立优化问题,在保证无偏梯度估计的前提下最小化残差误差,并显式考虑各计算节点的慢节点概率。通过拉格朗日对偶与凸优化推导出编码和解码系数的闭式解,同时提出数据分配策略,降低冗余与计算负载。我们还分析了在 $λ$-强凸与 $μ$-光滑损失函数下的收敛性。数值结果表明,该方法显著减轻慢节点影响,加速收敛,优于现有方法。

原文摘要 · Abstract (English)

In this paper, we propose an optimally structured gradient coding scheme to mitigate the straggler problem in distributed learning. Conventional gradient coding methods often assume homogeneous straggler models or rely on excessive data replication, limiting performance in real-world heterogeneous systems. To address these limitations, we formulate an optimization problem minimizing residual error while ensuring unbiased gradient estimation by explicitly considering individual straggler probabilities. We derive closed-form solutions for optimal encoding and decoding coefficients via Lagrangian duality and convex optimization, and propose data allocation strategies that reduce both redundancy and computation load. We also analyze convergence behavior for $λ$-strongly convex and $μ$-smooth loss functions. Numerical results show that our approach significantly reduces the impact of stragglers and accelerates convergence compared to existing methods.

分布式学习梯度编码慢节点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。