arXiv:2507.07137cs.LGcs.CL2025-07被引 1

用语言模型自动检测扩散模型去忆效果,发现去忆会伤害相关概念。

Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge

  • 利用语言模型提取世界知识,识别去忆可能损伤的邻近概念
  • 发现去忆后模型对相关概念的生成能力显著下降
  • 可自动生成对抗性提示绕过去忆,适合评估与优化去忆方法

机器去忆(MU)是一种低成本清除基础扩散模型中不良信息(如生成概念、偏见或模式)的有前景方法。尽管去忆比重新训练模型成本低得多,但验证信息是否完全清除仍具挑战且耗时。此外,去忆可能损害模型对周边保留概念的性能,难以判断模型是否仍适用于部署。我们提出 autoeval-dmun,一种自动化工具,利用(视觉-)语言模型全面评估扩散模型的去忆效果。针对目标概念,该工具从语言模型中提取结构化世界知识,识别可能受去忆影响的邻近概念,并生成对抗性提示以绕过去忆。我们使用该工具评估主流去忆方法,发现:(1)语言模型能建立邻近概念的语义排序,与去忆损伤高度相关;(2)可有效通过合成对抗性提示绕过去忆。

原文摘要 · Abstract (English)

Machine unlearning (MU) is a promising cost-effective method to cleanse undesired information (generated concepts, biases, or patterns) from foundational diffusion models. While MU is orders of magnitude less costly than retraining a diffusion model without the undesired information, it can be challenging and labor-intensive to prove that the information has been fully removed from the model. Moreover, MU can damage diffusion model performance on surrounding concepts that one would like to retain, making it unclear if the diffusion model is still fit for deployment. We introduce autoeval-dmun, an automated tool which leverages (vision-) language models to thoroughly assess unlearning in diffusion models. Given a target concept, autoeval-dmun extracts structured, relevant world knowledge from the language model to identify nearby concepts which are likely damaged by unlearning and to circumvent unlearning with adversarial prompts. We use our automated tool to evaluate popular diffusion model unlearning methods, revealing that language models (1) impose semantic orderings of nearby concepts which correlate well with unlearning damage and (2) effectively circumvent unlearning with synthetic adversarial prompts.

扩散模型去忆评估语言模型自动化检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。