arXiv:2505.17013cs.LGcs.CV2025-05NeurIPS被引 15

探究扩散模型中概念擦除的彻底性与检测方法

When Are Concepts Erased From Diffusion Models?

  • 提出两种概念擦除机制:干扰引导过程或降低生成概率
  • 通过多类探测技术发现概念可能未被完全移除
  • 适合关注模型安全与可控生成的研究者阅读

在概念擦除任务中,模型被修改以选择性阻止生成特定概念。尽管新方法快速涌现,但其擦除彻底性仍不明确。本文提出扩散模型中擦除的两种机制:(i) 干扰模型内部引导过程,(ii) 降低目标概念的无条件生成概率,可能导致其完全消失。为评估概念是否真正被擦除,我们引入一套独立的探测技术:提供视觉上下文、修改扩散轨迹、应用分类器引导,以及分析替代生成结果。实验揭示了在非对抗性文本输入下评估擦除鲁棒性的价值,并强调对扩散模型擦除效果进行综合评估的重要性。

原文摘要 · Abstract (English)

In concept erasure, a model is modified to selectively prevent it from generating a target concept. Despite the rapid development of new methods, it remains unclear how thoroughly these approaches remove the target concept from the model. We begin by proposing two conceptual models for the erasure mechanism in diffusion models: (i) interfering with the model's internal guidance processes, and (ii) reducing the unconditional likelihood of generating the target concept, potentially removing it entirely. To assess whether a concept has been truly erased from the model, we introduce a comprehensive suite of independent probing techniques: supplying visual context, modifying the diffusion trajectory, applying classifier guidance, and analyzing the model's alternative generations that emerge in place of the erased concept. Our results shed light on the value of exploring concept erasure robustness outside of adversarial text inputs, and emphasize the importance of comprehensive evaluations for erasure in diffusion models.

扩散模型概念擦除模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。