arXiv:2501.09833cs.CV2025-01ICCV被引 13

概念擦除会误伤无关概念,导致生成质量下降

Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts

  • 提出ErasureBench基准,系统评估擦除后模型表现
  • 发现擦除非目标概念时会产生连带抑制效应
  • 适合关注模型安全与可控生成的研究者参考

概念擦除技术近年来受到广泛关注,旨在从文生图模型中移除不想要的概念。尽管在控制环境下表现良好,其在真实场景中的鲁棒性与部署可行性仍不确定。本文(1)指出当前对净化模型评估的不足,尤其缺乏对多样化概念维度的全面考察;(2)系统分析了文生图模型在概念擦除后的失效模式,重点关注视觉相似、二元对立及语义相关等不同层级关系下非目标概念的意外影响。为此,我们构建了ErasureBench,一个涵盖100多个精选概念、定制化评估提示和一套稳健指标的综合性基准,用于评估擦除效果与副作用。结果揭示了一种概念纠缠现象:擦除操作会导致非目标概念的非预期抑制,引发扩散式退化,表现为生成畸变和质量下降。

原文摘要 · Abstract (English)

Concept erasure techniques have recently gained significant attention for their potential to remove unwanted concepts from text-to-image models. While these methods often demonstrate promising results in controlled settings, their robustness in real-world applications and suitability for deployment remain uncertain. In this work, we (1) identify a critical gap in evaluating sanitized models, particularly in assessing their performance across diverse concept dimensions, and (2) systematically analyze the failure modes of text-to-image models post-erasure. We focus on the unintended consequences of concept removal on non-target concepts across different levels of interconnected relationships including visually similar, binomial, and semantically related concepts. To address this, we introduce EraseBench, a comprehensive benchmark for evaluating post-erasure performance. EraseBench includes over 100 curated concepts, targeted evaluation prompts, and a robust set of metrics to assess both effectiveness and side effects of erasure. Our findings reveal a phenomenon of concept entanglement, where erasure leads to unintended suppression of non-target concepts, causing spillover degradation that manifests as distortions and a decline in generation quality.

概念擦除生成质量模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。