对比两种先进聚类联邦学习方法在不同数据异构场景下的表现
Comparative Evaluation of Clustered Federated Learning Methods
- 提出数据异构性的系统分类法,用于评估聚类联邦学习效果
- 在三个图像数据集上验证算法性能,发现异构类型影响聚类质量
- 为实际应用中选择合适算法提供可参考的实证依据
近年来,联邦学习(FL)作为保护数据隐私的分布式学习方法备受关注。随着其发展,面对真实场景中的新挑战,尤其是客户端间存在高度非独立同分布(non-IID)数据的问题日益凸显。聚类联邦学习(CFL)通过将客户端分组以实现组内数据分布一致,成为应对该问题的主流方案。然而,现有先进CFL算法通常仅在少数异构情形下测试,缺乏系统性验证,且异构场景的分类标准不清晰。本文针对联邦学习中的数据异构性提出一套系统分类体系,评估两种前沿CFL算法在三个图像分类数据集上的表现,并利用外部聚类指标分析结果与异构类别之间的关系。目标是厘清CFL性能与数据异构类型之间的关联,为实际部署提供更可靠的决策支持。
原文摘要 · Abstract (English)
Over recent years, Federated Learning (FL) has proven to be one of the most promising methods of distributed learning which preserves data privacy. As the method evolved and was confronted to various real-world scenarios, new challenges have emerged. One such challenge is the presence of highly heterogeneous (often referred as non-IID) data distributions among participants of the FL protocol. A popular solution to this hurdle is Clustered Federated Learning (CFL), which aims to partition clients into groups where the distribution are homogeneous. In the literature, state-of-the-art CFL algorithms are often tested using a few cases of data heterogeneities, without systematically justifying the choices. Further, the taxonomy used for differentiating the different heterogeneity scenarios is not always straightforward. In this paper, we explore the performance of two state-of-theart CFL algorithms with respect to a proposed taxonomy of data heterogeneities in federated learning (FL). We work with three image classification datasets and analyze the resulting clusters against the heterogeneity classes using extrinsic clustering metrics. Our objective is to provide a clearer understanding of the relationship between CFL performances and data heterogeneity scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。