arXiv:2504.01921cs.LGstat.ML2025-04被引 6

同时应对数据差异与网络延迟,提升联邦学习效率。

Client Selection in Federated Learning with Data Heterogeneity and Network Latencies

  • 基于理论最优目标,每轮优化选择客户端以减少收敛时间。
  • 在9个非独立同分布数据集上,性能最优比基线快20倍。
  • 适用于真实场景中数据与网络条件双重异构的联邦学习。

联邦学习(FL)是一种分布式机器学习范式,多个客户端基于私有数据进行本地训练,并将更新后的模型发送至中心服务器进行全局聚合。实际收敛受多种因素挑战,主要障碍是客户端间的异质性,表现为本地数据分布差异和模型传输过程中的延迟差异。现有研究虽提出多种高效客户端选择方法以缓解单一异质性的影响,但尚未有方法能有效应对两者同时存在的真实场景。本文提出两种理论上最优的客户端选择方案,可同时处理数据与延迟异质性。方法通过每轮求解简化优化问题,最小化理论收敛时间。在9个具有非独立同分布数据的基准、2种实际延迟分布及非凸神经网络模型上的实验表明,所提算法至少与最优基线相当,最多优于基线20倍。

原文摘要 · Abstract (English)

Federated learning (FL) is a distributed machine learning paradigm where multiple clients conduct local training based on their private data, then the updated models are sent to a central server for global aggregation. The practical convergence of FL is challenged by multiple factors, with the primary hurdle being the heterogeneity among clients. This heterogeneity manifests as data heterogeneity concerning local data distribution and latency heterogeneity during model transmission to the server. While prior research has introduced various efficient client selection methods to alleviate the negative impacts of either of these heterogeneities individually, efficient methods to handle real-world settings where both these heterogeneities exist simultaneously do not exist. In this paper, we propose two novel theoretically optimal client selection schemes that can handle both these heterogeneities. Our methods involve solving simple optimization problems every round obtained by minimizing the theoretical runtime to convergence. Empirical evaluations on 9 datasets with non-iid data distributions, 2 practical delay distributions, and non-convex neural network models demonstrate that our algorithms are at least competitive to and at most 20 times better than best existing baselines.

联邦学习客户端选择异质性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。