arXiv:2510.03380cs.LGcs.AI2025-10被引 3

提出新方法应对联邦学习中数据量差异大的难题

A Robust Clustered Federated Learning Approach for Non-IID Data with Quantity Skew

  • 融合客户端与服务器两种聚类策略,动态优化分组
  • 在270种非独立同分布配置下,准确率与聚类质量均最优
  • 适合数据量差异大、隐私要求高的实际应用

联邦学习(FL)是一种去中心化范式,允许客户端与服务器协作训练全局人工智能模型,而无需共享原始数据,从而保护隐私。其主要挑战之一是非独立同分布(Non-IID)数据,其中数据量偏斜(QS)尤为突出,即客户端持有的数据量高度异质。聚类联邦学习(CFL)是解决该问题的新兴方法,通过将具有相似数据分布的客户端分组以提升模型性能。现有CFL方法通常采用两类策略:一为客户端根据本地训练损失最小选择集群;二为服务器基于本地模型相似性进行分组。然而,多数CFL方法在数量偏斜(QS)场景下缺乏系统评估,且面临严峻挑战。本文提出两项贡献:一是对先进CFL算法在多种非独立同分布设置下的全面评估,涵盖多个数量偏斜场景;二是提出一种新型迭代式CFL算法CORNFLQS,实现两类策略间的最优协调。实验在六个图像分类数据集上展开,共生成270个非独立同分布配置。结果表明,CORNFLQS在准确率和聚类质量上均取得最高平均排名,并对各类数量偏斜扰动表现出强鲁棒性,整体优于现有CFL方法。

原文摘要 · Abstract (English)

Federated Learning (FL) is a decentralized paradigm that enables a client-server architecture to collaboratively train a global Artificial Intelligence model without sharing raw data, thereby preserving privacy. A key challenge in FL is Non-IID data. Quantity Skew (QS) is a particular problem of Non-IID, where clients hold highly heterogeneous data volumes. Clustered Federated Learning (CFL) is an emergent variant of FL that presents a promising solution to Non-IID problem. It improves models' performance by grouping clients with similar data distributions into clusters. CFL methods generally fall into two operating strategies. In the first strategy, clients select the cluster that minimizes the local training loss. In the second strategy, the server groups clients based on local model similarities. However, most CFL methods lack systematic evaluation under QS but present significant challenges because of it. In this paper, we present two main contributions. The first one is an evaluation of state-of-the-art CFL algorithms under various Non-IID settings, applying multiple QS scenarios to assess their robustness. Our second contribution is a novel iterative CFL algorithm, named CORNFLQS, which proposes an optimal coordination between both operating strategies of CFL. Our approach is robust against the different variations of QS settings. We conducted intensive experiments on six image classification datasets, resulting in 270 Non-IID configurations. The results show that CORNFLQS achieves the highest average ranking in both accuracy and clustering quality, as well as strong robustness to QS perturbations. Overall, our approach outperforms actual CFL algorithms.

联邦学习数据偏斜聚类隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。