揭穿扩散模型遗忘技术假象:看似删除实则隐藏概念
Unlearning or Concealment? A Critical Analysis and Evaluation Metrics for Unlearning in Diffusion Models
- 用对抗攻击检测遗忘后仍可恢复目标概念
- 提出两项新指标,量化概念残留程度
- 揭示当前主流方法实际是隐藏而非真正遗忘
近期研究对文本到图像扩散模型中的概念移除与定向遗忘方法表现出浓厚兴趣。本文通过全面的白盒分析,揭示了现有扩散模型遗忘方法的脆弱性。结果显示,现有方法导致目标概念与对应提示解耦,这实际上是概念隐藏而非真正的遗忘。本文对五种常用的扩散模型遗忘技术进行了严谨的理论与实证评估,揭示其潜在缺陷。我们提出两项新评价指标:概念重获得分(CRS)和概念置信得分(CCS)。CRS衡量遗忘后模型与全训练模型在潜在表示上的相似性,反映随引导强度增加时被遗忘概念的恢复程度;CCS量化模型对篡改数据分配目标概念的信心,反映未学习模型生成结果与原始知识对齐的概率。这两项指标使概念擦除效果评估更可靠。使用新指标评估五种顶尖方法,发现其真正遗忘能力存在显著不足。
原文摘要 · Abstract (English)
Recent research has seen significant interest in methods for concept removal and targeted forgetting in text-to-image diffusion models. In this paper, we conduct a comprehensive white-box analysis showing the vulnerabilities in existing diffusion model unlearning methods. We show that existing unlearning methods lead to decoupling of the targeted concepts (meant to be forgotten) for the corresponding prompts. This is concealment and not actual forgetting, which was the original goal. This paper presents a rigorous theoretical and empirical examination of five commonly used techniques for unlearning in diffusion models, while showing their potential weaknesses. We introduce two new evaluation metrics: Concept Retrieval Score (\textbf{CRS}) and Concept Confidence Score (\textbf{CCS}). These metrics are based on a successful adversarial attack setup that can recover \textit{forgotten} concepts from unlearned diffusion models. \textbf{CRS} measures the similarity between the latent representations of the unlearned and fully trained models after unlearning. It reports the extent of retrieval of the \textit{forgotten} concepts with increasing amount of guidance. CCS quantifies the confidence of the model in assigning the target concept to the manipulated data. It reports the probability of the \textit{unlearned} model's generations to be aligned with the original domain knowledge with increasing amount of guidance. The \textbf{CCS} and \textbf{CRS} enable a more robust evaluation of concept erasure methods. Evaluating existing five state-of-the-art methods with our metrics, reveal significant shortcomings in their ability to truly \textit{unlearn}. Source Code: \color{blue}{https://respailab.github.io/unlearning-or-concealment}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。