系统梳理联邦学习中非独立同分布数据的分类与应对方法
Non-IID data in Federated Learning: A Survey with Taxonomy, Metrics, Methods, Frameworks and Future Directions
- 构建非独立同分布数据的分类体系与量化指标
- 总结主流方法与框架在异构数据下的应用效果
- 适合关注联邦学习实际落地的研究者与工程师
近年来,机器学习的发展使联邦学习(FL)成为一种有前景的方法,允许多个分布式用户(即客户端)在不共享私有数据的情况下协同训练模型。尽管该隐私保护方法潜力巨大,但在客户端数据非独立且非同分布(non-IID)时仍面临挑战,导致模型性能下降和训练速度变慢。由于对non-IID数据的分类与量化缺乏共识,本技术综述旨在填补这一空白,提出详细的non-IID数据分类体系、数据划分协议及量化异质性的指标。同时,概述了应对non-IID数据的常用解决方案以及在异构数据下常用的标准化框架。基于最新研究,本文总结关键经验并提出未来研究方向。
原文摘要 · Abstract (English)
Recent advances in machine learning have highlighted Federated Learning (FL) as a promising approach that enables multiple distributed users (so-called clients) to collectively train ML models without sharing their private data. While this privacy-preserving method shows potential, it struggles when data across clients is not independent and identically distributed (non-IID) data. The latter remains an unsolved challenge that can result in poorer model performance and slower training times. Despite the significance of non-IID data in FL, there is a lack of consensus among researchers about its classification and quantification. This technical survey aims to fill that gap by providing a detailed taxonomy for non-IID data, partition protocols, and metrics to quantify data heterogeneity. Additionally, we describe popular solutions to address non-IID data and standardized frameworks employed in FL with heterogeneous data. Based on our state-of-the-art survey, we present key lessons learned and suggest promising future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。