arXiv:2410.04810cs.LGcs.CV2024-10CVPR被引 26

用个性化扩散模型解决异构数据下的单轮联邦学习难题

FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models

  • 双层次个性化扩散模型,分别适应客户端数据分布和概念特征
  • 在医疗与卫星图像上实现显著优于现有方法的性能提升
  • 适合数据稀疏、分布异构的现实联邦学习场景

单轮联邦学习(OSFL)作为一种去中心化机器学习范式,仅需一轮客户端数据或模型上传,有效降低通信开销并缓解隐私风险。然而,真实系统中客户端数据异构性和数据量有限导致现有方法性能受限。近期,潜在扩散模型(LDM)通过大规模预训练在高质量图像生成方面取得突破,为解决该问题提供可能。但直接应用预训练LDM会导致合成数据分布偏移,尤其在医学影像等低频领域表现更差。为此,本文提出联邦双层个性化(FedBiP),在实例级与概念级对预训练LDM进行个性化,生成符合客户端本地数据分布的图像,同时满足隐私要求。FedBiP是首个同时应对特征空间异构与数据稀缺问题的OSFL方法。在三个含特征异构的基准数据集及具有标签异构的医疗与卫星图像数据集上的实验表明,该方法显著优于其他OSFL方法。

原文摘要 · Abstract (English)

One-Shot Federated Learning (OSFL), a special decentralized machine learning paradigm, has recently gained significant attention. OSFL requires only a single round of client data or model upload, which reduces communication costs and mitigates privacy threats compared to traditional FL. Despite these promising prospects, existing methods face challenges due to client data heterogeneity and limited data quantity when applied to real-world OSFL systems. Recently, Latent Diffusion Models (LDM) have shown remarkable advancements in synthesizing high-quality images through pretraining on large-scale datasets, thereby presenting a potential solution to overcome these issues. However, directly applying pretrained LDM to heterogeneous OSFL results in significant distribution shifts in synthetic data, leading to performance degradation in classification models trained on such data. This issue is particularly pronounced in rare domains, such as medical imaging, which are underrepresented in LDM's pretraining data. To address this challenge, we propose Federated Bi-Level Personalization (FedBiP), which personalizes the pretrained LDM at both instance-level and concept-level. Hereby, FedBiP synthesizes images following the client's local data distribution without compromising the privacy regulations. FedBiP is also the first approach to simultaneously address feature space heterogeneity and client data scarcity in OSFL. Our method is validated through extensive experiments on three OSFL benchmarks with feature space heterogeneity, as well as on challenging medical and satellite image datasets with label heterogeneity. The results demonstrate the effectiveness of FedBiP, which substantially outperforms other OSFL methods.

联邦学习扩散模型数据异构个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。