提出新方法从扩散模型中提取训练数据,揭示无条件模型也存在隐私风险。
SIDE: Surrogate Conditional Data Extraction from Diffusion Models
- 用数据驱动的替代条件,实现对任意扩散模型的定向数据提取。
- 在多个数据集上成功提取无条件模型中的训练数据,效果优于传统攻击方法。
- 揭示所有形式的条件都会增强记忆性,为模型隐私评估提供新标准。
随着扩散概率模型(DPMs)在生成式人工智能中的核心地位日益凸显,理解其记忆行为对于评估数据泄露、版权侵权及可信度等风险至关重要。以往研究发现,条件化DPMs极易受到显式提示下的数据提取攻击,而无条件模型常被认为安全。本文挑战这一观点,提出 extbf{Surrogate conditional Data Extraction (SIDE)}——一种通用框架,通过构建数据驱动的替代条件,使任何DPM都能实现目标数据提取。在CIFAR-10、CelebA、ImageNet和LAION-5B上的大量实验表明,SIDE可成功从所谓‘安全’的无条件模型中提取训练数据,甚至优于现有基线攻击方法。此外,我们基于信息标签构建统一理论框架,证明所有形式的条件(显式或替代)均会放大模型的记忆能力。本工作重新定义了DPMs的威胁图景,确立精确条件作为根本性漏洞,并为模型隐私评估设定了更强的新基准。
原文摘要 · Abstract (English)
As diffusion probabilistic models (DPMs) become central to Generative AI (GenAI), understanding their memorization behavior is essential for evaluating risks such as data leakage, copyright infringement, and trustworthiness. While prior research finds conditional DPMs highly susceptible to data extraction attacks using explicit prompts, unconditional models are often assumed to be safe. We challenge this view by introducing \textbf{Surrogate condItional Data Extraction (SIDE)}, a general framework that constructs data-driven surrogate conditions to enable targeted extraction from any DPM. Through extensive experiments on CIFAR-10, CelebA, ImageNet, and LAION-5B, we show that SIDE can successfully extract training data from so-called safe unconditional models, outperforming baseline attacks even on conditional models. Complementing these findings, we present a unified theoretical framework based on informative labels, demonstrating that all forms of conditioning, explicit or surrogate, amplify memorization. Our work redefines the threat landscape for DPMs, establishing precise conditioning as a fundamental vulnerability and setting a new, stronger benchmark for model privacy evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。