arXiv:2508.15124cs.LGcs.CV2025-08EMNLP被引 9

发现扩散模型概念删除有副作用,易被绕过。

Side Effects of Erasing Concepts from Diffusion Models

  • 用层级组合提示词测试概念删除鲁棒性
  • 删除后仍能生成目标概念的变体和属性
  • 适合关注生成模型安全与隐私的研究者

文本到图像生成模型引发隐私、版权和安全担忧,催生了概念擦除技术(CETs)。理想状态下,CET应禁止生成用户指定的“目标”概念,同时保持对其他概念的高质量生成能力。本文揭示概念擦除存在副作用,且可轻易被绕过。为全面评估CET鲁棒性,我们提出侧效应评估(SEE)基准,包含描述物体及其属性的层级与组合式提示。该数据集与自动化评估流程量化了三方面影响:邻近概念受影响程度、目标概念规避情况、属性泄露问题。实验表明,通过上位-下位类别关系、语义相似提示及目标概念的组合变体,可轻松绕过CET。我们还发现CET存在属性泄露,以及注意力集中或分散的反直觉现象。我们公开了基准与评估工具,以支持未来鲁棒概念擦除研究。

原文摘要 · Abstract (English)

Concerns about text-to-image (T2I) generative models infringing on privacy, copyright, and safety have led to the development of concept erasure techniques (CETs). The goal of an effective CET is to prohibit the generation of undesired "target" concepts specified by the user, while preserving the ability to synthesize high-quality images of other concepts. In this work, we demonstrate that concept erasure has side effects and CETs can be easily circumvented. For a comprehensive measurement of the robustness of CETs, we present the Side Effect Evaluation (SEE) benchmark that consists of hierarchical and compositional prompts describing objects and their attributes. The dataset and an automated evaluation pipeline quantify side effects of CETs across three aspects: impact on neighboring concepts, evasion of targets, and attribute leakage. Our experiments reveal that CETs can be circumvented by using superclass-subclass hierarchy, semantically similar prompts, and compositional variants of the target. We show that CETs suffer from attribute leakage and a counterintuitive phenomenon of attention concentration or dispersal. We release our benchmark and evaluation tools to aid future work on robust concept erasure.

扩散模型概念删除安全性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。