arXiv:2503.04052cs.LG2025-03被引 1

研究异步联邦学习中延迟对模型性能的影响,发现丢弃旧数据未必最优。

The Impact Analysis of Delays in Asynchronous Federated Learning with Data Heterogeneity for Edge Intelligence

  • 提出异步误差定义,分析数据异构在同步框架下的负面影响。
  • 发现客户端平均延迟越短,训练效果不一定越好。
  • 提出重用延迟梯度的伪同步方案,理论证明可提升收敛性。

联邦学习(FL)为边缘智能中的多设备协同训练提供了新范式,但实际应用中面临数据异构和通信计算延迟等挑战。本文研究异步联邦学习(AFL)系统中未知延迟对训练性能的影响。首先,在传统同步联邦学习(SFL)框架下,基于新提出的异步误差定义,理论上分析了数据异构的单一负面影响。接着,分析了常规的异步更新延迟梯度(AUDG)方案,发现数据异构的负面效应与延迟相关,且特定客户端平均延迟较短并不总能提升性能。为弥补AUDG不适用场景,提出伪同步更新重用延迟梯度(PSURDG)方案,并给出理论收敛性分析。在两者中,每轮仅随机部分客户端成功上传更新结果,核心差异在于是否重用延迟信息。通过理论分析与仿真验证,直观表明因时间延迟而丢弃过时信息并非总是最佳策略。

原文摘要 · Abstract (English)

Federated learning (FL) has provided a new methodology for coordinating a group of clients to train a machine learning model collaboratively, bringing an efficient paradigm in edge intelligence. Despite its promise, FL faces several critical challenges in practical applications involving edge devices, such as data heterogeneity and delays stemming from communication and computation constraints. This paper examines the impact of unknown causes of delay on training performance in an Asynchronous Federated Learning (AFL) system with data heterogeneity. Initially, an asynchronous error definition is proposed, based on which the solely adverse impact of data heterogeneity is theoretically analyzed within the traditional Synchronous Federated Learning (SFL) framework. Furthermore, Asynchronous Updates with Delayed Gradients (AUDG), a conventional AFL scheme, is discussed. Investigation into AUDG reveals that the negative influence of data heterogeneity is correlated with delays, while a shorter average delay from a specific client does not consistently enhance training performance. In order to compensate for the scenarios where AUDG are not adapted, Pseudo-synchronous Updates by Reusing Delayed Gradients (PSURDG) is proposed, and its theoretical convergence is analyzed. In both AUDG and PSURDG, only a random set of clients successfully transmits their updated results to the central server in each iteration. The critical difference between them lies in whether the delayed information is reused. Finally, both schemes are validated and compared through theoretical analysis and simulations, demonstrating more intuitively that discarding outdated information due to time delays is not always the best approach.

联邦学习异步训练边缘智能数据异构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。