厘清聚类联邦学习的分类与应用,区分隐私与效率的权衡。
A survey on Clustered Federated Learning: Taxonomy, Analysis and Applications
- 按服务器、客户端、元数据三类划分聚类联邦学习方法。
- 理论研究重隐私,实际应用更关注效率与系统异构性。
- 明确区分核心聚类联邦学习与操作型变体,避免概念混淆。
随着联邦学习(FL)的发展,非独立同分布(non-IID)数据问题日益突出。聚类联邦学习(CFL)通过为具有相似数据分布的客户端群组训练多个专用模型来应对这一挑战。然而,'CFL'一词被广泛用于与数据异构性无关的操作策略,造成概念模糊。本文系统梳理了CFL文献,提出一个严谨的分类体系,将算法分为服务器端、客户端和基于元数据三类。分析显示:理论研究侧重于保护隐私的服务器/客户端方法,而物联网(IoT)、移动计算和能源等实际应用则普遍采用基于元数据的高效方案。此外,我们明确区分了‘核心CFL’(针对非IID数据分组)与‘聚类X FL’(用于系统异构性的操作变体)。最后,总结经验教训并提出未来方向,以弥合理论隐私与实践效率之间的鸿沟。
原文摘要 · Abstract (English)
As Federated Learning (FL) expands, the challenge of non-independent and identically distributed (non-IID) data becomes critical. Clustered Federated Learning (CFL) addresses this by training multiple specialized models, each representing a group of clients with similar data distributions. However, the term ''CFL'' has increasingly been applied to operational strategies unrelated to data heterogeneity, creating significant ambiguity. This survey provides a systematic review of the CFL literature and introduces a principled taxonomy that classifies algorithms into Server-side, Client-side, and Metadata-based approaches. Our analysis reveals a distinct dichotomy: while theoretical research prioritizes privacy-preserving Server/Client-side methods, real-world applications in IoT, Mobility, and Energy overwhelmingly favor Metadata-based efficiency. Furthermore, we explicitly distinguish ''Core CFL'' (grouping clients for non-IID data) from ''Clustered X FL'' (operational variants for system heterogeneity). Finally, we outline lessons learned and future directions to bridge the gap between theoretical privacy and practical efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。