arXiv:2504.05849cs.CV2025-04ICCV

用条件扩散模型做数据匿名化不安全,会泄露隐私。

On the Importance of Conditioning for Privacy-Preserving Data Augmentation

  • 用深度图等条件引导扩散生成图像
  • 攻击者可利用边缘特征识别原人物
  • 适合研究隐私保护与对抗攻击的读者

潜在扩散模型可用于增强训练数据,但若在生成时使用深度图或边缘图等条件,反而会暴露隐私。我们采用对比学习方法训练识别模型,能从候选池中准确找出目标人物。实验表明,基于条件扩散的匿名化易受黑盒攻击,原因在于生成图像保留了与原始图像相似的边缘特征,使识别模型可学习这些模式进行身份还原。

原文摘要 · Abstract (English)

Latent diffusion models can be used as a powerful augmentation method to artificially extend datasets for enhanced training. To the human eye, these augmented images look very different to the originals. Previous work has suggested to use this data augmentation technique for data anonymization. However, we show that latent diffusion models that are conditioned on features like depth maps or edges to guide the diffusion process are not suitable as a privacy preserving method. We use a contrastive learning approach to train a model that can correctly identify people out of a pool of candidates. Moreover, we demonstrate that anonymization using conditioned diffusion models is susceptible to black box attacks. We attribute the success of the described methods to the conditioning of the latent diffusion model in the anonymization process. The diffusion model is instructed to produce similar edges for the anonymized images. Hence, a model can learn to recognize these patterns for identification.

隐私保护扩散模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。