arXiv:2504.00952cs.LGcs.AI2025-04

用隐私保护的联邦学习训练扩散模型,生成更公平的合成数据。

Personalized Federated Training of Diffusion Models with Privacy Guarantees

  • 通过个性化联邦学习与扩散过程噪声,实现隐私安全的模型训练。
  • 在数据异构场景下性能优于非协作训练,减少合成数据偏差。
  • 适合医疗、金融等需高隐私保护的敏感领域应用。

在医疗、金融和生物医学研究等敏感领域,获取可访问、合规且合乎伦理的数据面临巨大挑战。同时,由于对隐私、版权和竞争的担忧,开放公共数据集的获取日益受限。合成数据成为有前景的替代方案,而扩散模型——一种前沿的生成式AI技术——能有效生成高质量且多样化的合成数据。本文提出一种新型联邦学习框架,用于在去中心化的私有数据集上训练扩散模型。该框架利用个性化机制和前向扩散过程中的固有噪声,在保证强差分隐私的前提下生成高质量样本。实验表明,该框架在数据高度异构的场景下优于非协作训练方法,有效降低合成数据中的偏差与不平衡,使下游模型更加公平。

原文摘要 · Abstract (English)

The scarcity of accessible, compliant, and ethically sourced data presents a considerable challenge to the adoption of artificial intelligence (AI) in sensitive fields like healthcare, finance, and biomedical research. Furthermore, access to unrestricted public datasets is increasingly constrained due to rising concerns over privacy, copyright, and competition. Synthetic data has emerged as a promising alternative, and diffusion models -- a cutting-edge generative AI technology -- provide an effective solution for generating high-quality and diverse synthetic data. In this paper, we introduce a novel federated learning framework for training diffusion models on decentralized private datasets. Our framework leverages personalization and the inherent noise in the forward diffusion process to produce high-quality samples while ensuring robust differential privacy guarantees. Our experiments show that our framework outperforms non-collaborative training methods, particularly in settings with high data heterogeneity, and effectively reduces biases and imbalances in synthetic data, resulting in fairer downstream models.

联邦学习扩散模型隐私保护合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。