arXiv:2506.17975cs.CV2025-06被引 2

用多样性感知扩散模型生成隐私安全的合成医学影像数据

Enabling PSO-Secure Synthetic Data Sharing Using Diversity-Aware Diffusion Models

  • 基于多样性感知的扩散模型训练框架,提升合成数据覆盖度
  • 合成数据在下游任务中性能仅比真实数据低1个百分点
  • 同时满足PSO隐私保护要求,适合医疗数据共享场景

合成数据已达到与真实数据难以区分的视觉保真度,为医学影像中的隐私保护数据共享提供了巨大潜力。然而,完全合成的数据集仍存在显著局限:首先,数据共享的法律合规性常被忽视,如GDPR等法规未被充分遵守;其次,合成模型在域内下游任务中的性能仍不及真实数据。近期图像生成方法更关注最大化图像多样性以提升模式覆盖和下游性能。本文提出新视角:最大化多样性可有效防止自然人被单独识别,从而实现谓词单点识别(PSO)安全的合成数据。我们构建了一个通用框架,在个人数据上训练扩散模型,生成去个性化合成数据,在下游任务中性能仅比真实数据模型低1个百分点,显著优于不保证隐私的现有方法。代码已开源:https://github.com/MischaD/Trichotomy。

原文摘要 · Abstract (English)

Synthetic data has recently reached a level of visual fidelity that makes it nearly indistinguishable from real data, offering great promise for privacy-preserving data sharing in medical imaging. However, fully synthetic datasets still suffer from significant limitations: First and foremost, the legal aspect of sharing synthetic data is often neglected and data regulations, such as the GDPR, are largley ignored. Secondly, synthetic models fall short of matching the performance of real data, even for in-domain downstream applications. Recent methods for image generation have focused on maximising image diversity instead of fidelity solely to improve the mode coverage and therefore the downstream performance of synthetic data. In this work, we shift perspective and highlight how maximizing diversity can also be interpreted as protecting natural persons from being singled out, which leads to predicate singling-out (PSO) secure synthetic datasets. Specifically, we propose a generalisable framework for training diffusion models on personal data which leads to unpersonal synthetic datasets achieving performance within one percentage point of real-data models while significantly outperforming state-of-the-art methods that do not ensure privacy. Our code is available at https://github.com/MischaD/Trichotomy.

合成数据扩散模型隐私保护医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。