arXiv:2512.22877cs.CV2025-12被引 2

首个评估扩散模型多模态概念擦除的基准,发现现有方法在图像编辑中极易失效。

M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion Models

  • 构建文本、嵌入、反演隐变量三模态评估框架,覆盖白盒与黑盒场景
  • 90%以上概念在反演隐变量下复现,暴露现有方法严重漏洞
  • 提出IRECE模块,推理时增强鲁棒性,最多降低40%复现率

文本到图像扩散模型可能生成有害或受版权保护的内容,促使概念擦除研究发展。然而,现有方法主要关注文本提示中的概念擦除,忽视了图像编辑和个性化生成中日益关键的其他输入模态。这些模态可能成为攻击面,导致擦除后概念仍会重现。为此,我们引入M-ErasureBench,一个新型多模态评估框架,系统评测概念擦除方法在三种输入模态下的表现:文本提示、学习嵌入和反演隐变量。针对后两者,我们评估白盒与黑盒访问,共形成五种评估场景。分析显示,现有方法在文本提示上表现良好,但在学习嵌入和反演隐变量下普遍失效,白盒设置下概念复现率(CRR)超过90%。为应对这些漏洞,我们提出IRECE(推理时鲁棒性增强的概念擦除),一个即插即用模块,通过交叉注意力定位目标概念,并在去噪过程中扰动相关隐变量。实验表明,IRECE在最具挑战性的白盒隐变量反演场景下,将CRR最多降低40%,同时保持视觉质量。据我们所知,M-ErasureBench是首个超越文本提示的概念擦除综合基准。结合IRECE,该基准为构建更可靠的保护性生成模型提供了实用保障。

原文摘要 · Abstract (English)

Text-to-image diffusion models may generate harmful or copyrighted content, motivating research on concept erasure. However, existing approaches primarily focus on erasing concepts from text prompts, overlooking other input modalities that are increasingly critical in real-world applications such as image editing and personalized generation. These modalities can become attack surfaces, where erased concepts re-emerge despite defenses. To bridge this gap, we introduce M-ErasureBench, a novel multimodal evaluation framework that systematically benchmarks concept erasure methods across three input modalities: text prompts, learned embeddings, and inverted latents. For the latter two, we evaluate both white-box and black-box access, yielding five evaluation scenarios. Our analysis shows that existing methods achieve strong erasure performance against text prompts but largely fail under learned embeddings and inverted latents, with Concept Reproduction Rate (CRR) exceeding 90% in the white-box setting. To address these vulnerabilities, we propose IRECE (Inference-time Robustness Enhancement for Concept Erasure), a plug-and-play module that localizes target concepts via cross-attention and perturbs the associated latents during denoising. Experiments demonstrate that IRECE consistently restores robustness, reducing CRR by up to 40% under the most challenging white-box latent inversion scenario, while preserving visual quality. To the best of our knowledge, M-ErasureBench provides the first comprehensive benchmark of concept erasure beyond text prompts. Together with IRECE, our benchmark offers practical safeguards for building more reliable protective generative models.

扩散模型概念擦除多模态安全生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。