arXiv:2501.02219cs.LGcs.AI2025-01中稿 · IEEE WCNC 2025

用扩散模型生成数据,提升联邦半监督学习的准确率

Diffusion Model-Based Data Synthesis Aided Federated Semi-Supervised Learning

  • 用扩散模型合成缺失类别的数据,缓解标签数据不足问题
  • 在CIFAR-10上仅10%标签数据时,准确率从38.46%提升至52.14%
  • 适合标签稀少且数据分布异质的联邦学习场景

联邦半监督学习(FSSL)主要面临两个挑战:客户端标签数据稀缺,以及客户端间数据分布非独立同分布(non-IID)。本文提出一种新方法——基于扩散模型的数据合成辅助联邦半监督学习(DDSA-FSSL),利用扩散模型(DM)生成合成数据,弥合局部数据分布与全局分布之间的差异。在DDSA-FSSL中,客户端通过联邦学习训练的分类器对无标签数据进行伪标注,再以标注数据和优化后的伪标签数据协同训练扩散模型,从而生成其本地标签数据集中缺失类别的合成样本。该过程使客户端生成更全面、符合全局分布的合成数据集。在多个数据集及不同non-IID设置下的大量实验表明,该方法有效,例如在仅10%标签数据的CIFAR-10上,准确率从38.46%提升至52.14%。

原文摘要 · Abstract (English)

Federated semi-supervised learning (FSSL) is primarily challenged by two factors: the scarcity of labeled data across clients and the non-independent and identically distribution (non-IID) nature of data among clients. In this paper, we propose a novel approach, diffusion model-based data synthesis aided FSSL (DDSA-FSSL), which utilizes a diffusion model (DM) to generate synthetic data, bridging the gap between heterogeneous local data distributions and the global data distribution. In DDSA-FSSL, clients address the challenge of the scarcity of labeled data by employing a federated learning-trained classifier to perform pseudo labeling for unlabeled data. The DM is then collaboratively trained using both labeled and precision-optimized pseudo-labeled data, enabling clients to generate synthetic samples for classes that are absent in their labeled datasets. This process allows clients to generate more comprehensive synthetic datasets aligned with the global distribution. Extensive experiments conducted on multiple datasets and varying non-IID distributions demonstrate the effectiveness of DDSA-FSSL, e.g., it improves accuracy from 38.46% to 52.14% on CIFAR-10 datasets with 10% labeled data.

联邦学习扩散模型半监督数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。