arXiv:2507.16880cs.CVcs.AI2025-07

发现扩散模型中记忆图像的隐藏触发机制,揭示现有清除方法的脆弱性。

Finding DoRI: Discovery of Retained Images in Diffusion Models

  • 通过扰动文本嵌入重现被清除的训练图像,验证记忆非局部特性。
  • 相同图像在不同嵌入下激活模式差异大,表明记忆分布广泛。
  • 提出对抗微调新方法,突破局部假设实现更鲁棒的隐私保护。

文生图扩散模型虽取得显著进展,但存在数据隐私与知识产权隐患,因其可能无意中记忆并复现训练数据。现有缓解方法基于记忆可定位的假设,通过识别并剪枝触发复现的权重来应对。我们挑战这一假设,证明即使经过剪枝,对已缓解提示的文本嵌入施加微小扰动仍可重新触发数据复现,显示此类方法的脆弱性。进一步分析表明,记忆并非固有局部:(1)复现触发点分布于整个文本嵌入空间;(2)生成同一图像的不同嵌入导致模型激活差异显著;(3)不同剪枝方法为同一图像识别出不一致的敏感权重集。最后,我们展示摆脱局部假设可通过对抗微调实现更稳健的缓解。这些发现深化了对文生图模型记忆本质的理解,并为未来更可靠的缓解方法提供指导。

原文摘要 · Abstract (English)

Text-to-image diffusion models (DMs) have achieved remarkable success in image generation. However, concerns about data privacy and intellectual property remain due to their potential to inadvertently memorize and replicate training data. Recent mitigation efforts have focused on identifying and pruning weights responsible for triggering verbatim training data replication, based on the assumption that memorization can be localized. We challenge this assumption and demonstrate that, even after such pruning, small perturbations to the text embeddings of previously mitigated prompts can re-trigger data replication, revealing the fragility of such methods. Our further analysis then provides multiple indications that memorization is indeed \textit{not} inherently local: (1) replication triggers for memorized images are distributed throughout text embedding space; (2) embeddings yielding the same replicated image produce divergent model activations; and (3) different pruning methods identify inconsistent sets of memorization-related weights for the same image. Finally, we show that bypassing the locality assumption enables more robust mitigation through adversarial fine-tuning. These findings provide new insights into the fundamental nature of memorization in text-to-image DMs and inform the future development of more reliable mitigation methods against DM memorization.

扩散模型记忆泄露隐私安全对抗训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。