arXiv:2604.01514cs.CLcs.CV2026-04

指令式遗忘在扩散模型中失效,因提示无法持续抑制目标概念

Why Instruction-Based Unlearning Fails in Diffusion Models?

  • 通过自然语言指令控制扩散模型生成过程以消除特定概念
  • 实验显示指令无法持续降低目标概念的注意力,导致概念仍被生成
  • 适合研究生成模型可解释性与安全控制的学者

指令式遗忘在大型语言模型推理时表现出良好效果,但其能否适用于其他生成模型尚不明确。本文研究了基于扩散的图像生成模型中的指令式遗忘,通过多概念和多种提示变体的受控实验,发现仅依赖自然语言遗忘指令时,扩散模型系统性地无法抑制目标概念。通过对去噪过程中CLIP文本编码器和交叉注意力动态的分析,我们发现遗忘指令未能引发对目标概念标记的持续注意力下降,导致目标概念表征在整个生成过程中持续存在。这些结果揭示了提示级指令在扩散模型中的根本局限性,表明有效遗忘需超越推理时的语言控制。

原文摘要 · Abstract (English)

Instruction-based unlearning has proven effective for modifying the behavior of large language models at inference time, but whether this paradigm extends to other generative models remains unclear. In this work, we investigate instruction-based unlearning in diffusion-based image generation models and show, through controlled experiments across multiple concepts and prompt variants, that diffusion models systematically fail to suppress targeted concepts when guided solely by natural-language unlearning instructions. By analyzing both the CLIP text encoder and cross-attention dynamics during the denoising process, we find that unlearning instructions do not induce sustained reductions in attention to the targeted concept tokens, causing the targeted concept representations to persist throughout generation. These results reveal a fundamental limitation of prompt-level instruction in diffusion models and suggest that effective unlearning requires interventions beyond inference-time language control.

扩散模型遗忘机制可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。