构建空间分布非独立同分布数据的联邦学习基准,更真实评估算法性能。
ProFed: a Benchmark for Proximity-based non-IID Federated Learning
- 基于地理区域模拟数据分布偏移,生成不同偏度的数据划分。
- 在MNIST、FashionMNIST等经典数据集上构建多场景非独立同分布数据集。
- 为联邦学习算法提供可复现、标准化的评估框架,适合研究者对比优化。
近年来,联邦学习(Federated Learning, FL)在机器学习领域受到广泛关注。尽管已有多种FL算法被提出,但当客户端间数据呈现非独立同分布(non-IID)时,其性能常显著下降。这种数据分布偏差常源于地理模式,例如文本数据中的区域语言差异或城市交通中的局部模式,导致特定区域内数据呈独立同分布,而跨区域则呈现非独立同分布。然而,现有算法评估通常通过随机方式分割非独立同分布数据,忽视了空间分布特性。为此,我们提出ProFed,一个模拟不同区域间数据偏度的基准,整合文献中多种偏度生成方法,并应用于MNIST、FashionMNIST、CIFAR-10和CIFAR-100等知名数据集。目标是为研究者提供一个标准化框架,以更有效、一致地评估联邦学习算法,并与现有基线进行比较。
原文摘要 · Abstract (English)
In recent years, cro:flFederated learning (FL) has gained significant attention within the machine learning community. Although various FL algorithms have been proposed in the literature, their performance often degrades when data across clients is non-independently and identically distributed (non-IID). This skewness in data distribution often emerges from geographic patterns, with notable examples including regional linguistic variations in text data or localized traffic patterns in urban environments. Such scenarios result in IID data within specific regions but non-IID data across regions. However, existing FL algorithms are typically evaluated by randomly splitting non-IID data across devices, disregarding their spatial distribution. To address this gap, we introduce ProFed, a benchmark that simulates data splits with varying degrees of skewness across different regions. We incorporate several skewness methods from the literature and apply them to well-known datasets, including MNIST, FashionMNIST, CIFAR-10, and CIFAR-100. Our goal is to provide researchers with a standardized framework to evaluate FL algorithms more effectively and consistently against established baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。