arXiv:2504.19903cs.LG2025-04

解析异步联邦学习中压缩、延迟与数据异构的相互影响

Analysis of Asynchronous Federated Learning: Unraveling the Interactions between Gradient Compression, Delay, and Data Heterogeneity

  • 建立更宽松假设下的收敛分析,提升异步联邦学习理论深度
  • 发现压缩率与延迟存在非线性交互,加剧误差传播
  • 证明误差反馈可有效抑制噪声,适合复杂真实场景

在实际联邦学习中,客户端与服务器间通信开销是主要瓶颈。梯度压缩可降低开销,误差反馈(EF)能恢复精度。由于设备异构,同步联邦学习常受慢节点拖累,而异步联邦学习可缓解此问题。然而,在异步设置下,异步延迟、数据异构和灵活参与这三重挑战与压缩/EF机制间的复杂交互仍缺乏理论理解。本文通过全面收敛分析,解耦并揭示这些交互关系。首先研究基础异步框架AsynFL,提出更优收敛性分析,收敛速率优于以往工作。接着扩展至压缩版本AsynFLC,推导其收敛充分条件,揭示延迟与压缩率间的非线性交互。分析表明,延迟与数据异构共同加剧压缩引入的误差,阻碍收敛。进一步研究集成EF的AsynFLC-EF,证明EF可有效降低梯度估计方差,在多重挑战下使收敛速率匹配原版AsynFL。且延迟与灵活参与仅影响高阶收敛项。实验结果验证了理论发现。

原文摘要 · Abstract (English)

In practical federated learning (FL), the large communication overhead between clients and the server is often a significant bottleneck. Gradient compression methods can effectively reduce this overhead, while error feedback (EF) restores model accuracy. Moreover, due to device heterogeneity, synchronous FL often suffers from stragglers and inefficiency-issues that asynchronous FL effectively alleviates. However, in asynchronous FL settings-which inherently face three major challenges: asynchronous delay, data heterogeneity, and flexible client participation-the complex interactions among these system/statistical constraints and compression/EF mechanisms remain poorly understood theoretically. In this paper, we fill this gap through a comprehensive convergence study that adequately decouples and unravels these complex interactions across various FL frameworks. We first consider a basic asynchronous FL framework AsynFL, and establish an improved convergence analysis that relies on fewer assumptions and yields a superior convergence rate than prior studies. We then extend our study to a compressed version, AsynFLC, and derive sufficient conditions for its convergence, indicating the nonlinear interaction between asynchronous delay and compression rate. Our analysis further demonstrates how asynchronous delay and data heterogeneity jointly exacerbate compression-induced errors, thereby hindering convergence. Furthermore, we study the convergence of AsynFLC-EF, the framework that further integrates EF. We prove that EF can effectively reduce the variance of gradient estimation under the aforementioned challenges, enabling AsynFLC-EF to match the convergence rate of AsynFL. We also show that the impact of asynchronous delay and flexible participation on EF is limited to slowing down the higher-order convergence term. Experimental results substantiate our analytical findings very well.

联邦学习异步训练梯度压缩误差反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。