arXiv:2502.13989cs.CVcs.AI2025-02被引 3

提出量化评估文本到图像模型概念擦除效果的新方法,避免依赖主观视觉判断。

Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models

  • 设计三重评估标准:目标概念响应、相关概念抑制、其他概念保留。
  • 实验发现现有方法在部分标准上表现不佳,存在明显缺陷。
  • 适合关注生成模型可控性与安全性的研究人员使用。

已有研究尝试从预训练的文生图模型中消除特定概念,但当前性能评估仍主要依赖可视化,结果易受主观判断影响。不同研究采用的量化指标各异,难以进行系统比较。本文提出EraseEval,一种新型评估方法,包含三个核心标准:(1)含目标概念的提示是否被正确响应;(2)与被擦除概念相关的概念是否有效抑制;(3)其他无关概念是否得以保留。这三个维度被整合为单一评分,若任一维度表现差则得分降低,实现更稳健的评估。我们对基线概念擦除方法进行了实验评估,梳理其特性并揭示其局限性。尽管这些标准看似基础,部分方法仍未能取得高分,指向未来研究方向。代码已开源于https://github.com/fmp453/erase-eval。

原文摘要 · Abstract (English)

Studies have been conducted to prevent specific concepts from being generated from pretrained text-to-image generative models, achieving concept erasure in various ways. However, the performance evaluation of these studies is still largely reliant on visualization, with the superiority of studies often determined by human subjectivity. The metrics of quantitative evaluation also vary, making comprehensive comparisons difficult. We propose EraseEval, an evaluation method that differs from previous evaluation methods in that it involves three fundamental evaluation criteria: (1) How well does the prompt containing the target concept be reflected, (2) To what extent the concepts related to the erased concept can reduce the impact of the erased concept, and (3) Whether other concepts are preserved. These criteria are evaluated and integrated into a single metric, such that a lower score is given if any of the evaluations are low, leading to a more robust assessment. We experimentally evaluated baseline concept erasure methods, organized their characteristics, and identified challenges with them. Despite being fundamental evaluation criteria, some concept erasure methods failed to achieve high scores, which point toward future research directions for concept erasure methods. Our code is available at https://github.com/fmp453/erase-eval.

概念擦除图像生成评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。