arXiv:2504.11216cs.LGcs.AI2025-04被引 1

提出新方法提升联邦学习中异构数据的训练效果

FedDiverse: Tackling Data Heterogeneity in Federated Learning with Diversity-Driven Client Selection

  • 用6个指标量化数据异构性,精准刻画分布差异
  • 构建7个涵盖多种异构场景的图像分类数据集
  • 通过互补客户端协作提升模型性能,适合真实场景应用

联邦学习(FL)可在保护隐私的前提下实现分布式数据上的模型训练。然而,在实际场景中,客户端数据常呈现非独立同分布且不平衡的问题,导致统计异构性,影响服务器模型在各客户端上的泛化能力,减缓收敛速度并降低性能。本文首先提出一种基于6项指标的统计异构性表征方法,涵盖全局与客户端属性不平衡、类别不平衡及伪相关性;随后构建并共享7个用于二分类和多分类图像识别任务的联邦学习数据集,覆盖广泛的真实世界异构场景;最后提出FEDDIVERSE算法,通过促进具有互补数据分布的客户端协作,有效管理并利用数据异构性。在所提7个数据集上的实验表明,该方法能显著提升多种联邦学习方法的性能与鲁棒性,同时保持低通信与计算开销。

原文摘要 · Abstract (English)

Federated Learning (FL) enables decentralized training of machine learning models on distributed data while preserving privacy. However, in real-world FL settings, client data is often non-identically distributed and imbalanced, resulting in statistical data heterogeneity which impacts the generalization capabilities of the server's model across clients, slows convergence and reduces performance. In this paper, we address this challenge by proposing first a characterization of statistical data heterogeneity by means of 6 metrics of global and client attribute imbalance, class imbalance, and spurious correlations. Next, we create and share 7 computer vision datasets for binary and multiclass image classification tasks in Federated Learning that cover a broad range of statistical data heterogeneity and hence simulate real-world situations. Finally, we propose FEDDIVERSE, a novel client selection algorithm in FL which is designed to manage and leverage data heterogeneity across clients by promoting collaboration between clients with complementary data distributions. Experiments on the seven proposed FL datasets demonstrate FEDDIVERSE's effectiveness in enhancing the performance and robustness of a variety of FL methods while having low communication and computational overhead.

联邦学习数据异构客户端选择图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。