arXiv:2503.20166cs.LGcs.DC2025-03

用AI生成内容缓解边缘联邦学习中的数据异构问题

AIGC-assisted Federated Learning for Edge Intelligence: Architecture Design, Research Challenges and Future Directions

  • 提出GenFL架构,利用扩散模型生成合成数据
  • 在CIFAR10/100上提升非独立同分布场景下的性能
  • 适合关注隐私计算与AI生成技术融合的研究者

联邦学习(FL)可在保障隐私安全的前提下利用大规模终端数据,是集中式机器学习的分布式替代方案。然而,数据异构性限制了其性能。为此,人工智能生成内容(AIGC)这一创新的数据合成技术成为潜在解决方案。本文首先概述了AIGC辅助联邦学习的系统架构、性能指标与挑战。随后提出生成式联邦学习(GenFL)架构,设计其工作流程及聚合与权重策略。基于CIFAR10与CIFAR100数据集,采用扩散模型生成数据以提升FL性能。在多种非独立同分布(non-IID)数据分布下进行实验,验证了GenFL在克服数据异构导致瓶颈方面的有效性。同时探讨了AIGC辅助联邦学习的开放研究方向。

原文摘要 · Abstract (English)

Federated learning (FL) can fully leverage large-scale terminal data while ensuring privacy and security, and is considered as a distributed alternative for the centralized machine learning. However, the issue of data heterogeneity poses limitations on FL's performance. To address this challenge, artificial intelligence-generated content (AIGC) which is an innovative data synthesis technique emerges as one potential solution. In this article, we first provide an overview of the system architecture, performance metrics, and challenges associated with AIGC-assistant FL system design. We then propose the Generative federated learning (GenFL) architecture and present its workflow, including the design of aggregation and weight policy. Finally, using the CIFAR10 and CIFAR100 datasets, we employ diffusion models to generate dataset and improve FL performance. Experiments conducted under various non-independent and identically distributed (non-IID) data distributions demonstrate the effectiveness of GenFL on overcoming the bottlenecks in FL caused by data heterogeneity. Open research directions in the research of AIGC-assisted FL are also discussed.

联邦学习AIGC边缘智能数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。