arXiv:2604.03114cs.CVcs.AI2026-04被引 1

测试视觉大模型能否真正遗忘敏感概念,发现提示抑制≠真实擦除。

Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning

  • 用提示和指令实现无训练遗忘,避免微调损伤模型能力。
  • 真实场景下提示效果有限,仅在已知目标时才能有效降低识别率。
  • 物体与场景概念最难抑制,强指令调优模型仍能识别被要求遗忘的内容。

基于网络规模数据训练的视觉语言模型(VLM)会保留敏感和受版权保护的视觉概念,部署时可能需要移除。基于微调的遗忘方法存在结构性缺陷:在狭窄的遗忘数据集上微调会损害模型通用能力,导致无法区分性能下降是否由遗忘过程引起。无训练方法通过提示或系统指令抑制概念,但缺乏严谨评估基准。本文提出 VLM-UnBench,首个针对 VLM 中无训练视觉概念遗忘的基准评测体系,涵盖四种遗忘程度、7 个源数据集、11 个概念轴,并结合三级探针分类法与五种评估条件,以区分真实遗忘与指令合规性。在 8 种评估设置和 13 种 VLM 配置下,现实中的遗忘提示使遗忘准确率接近无指令基线;仅在提供目标概念的“预言”条件下才出现显著下降。物体与场景概念最难以抑制,即使在明确指令下,强指令调优模型仍保持识别能力。结果揭示提示级抑制与真实视觉概念擦除之间存在明显差距。

原文摘要 · Abstract (English)

VLMs trained on web-scale data retain sensitive and copyrighted visual concepts that deployment may require removing. Training-based unlearning methods share a structural flaw: fine-tuning on a narrow forget set degrades general capabilities before unlearning begins, making it impossible to attribute subsequent performance drops to the unlearning procedure itself. Training-free approaches sidestep this by suppressing concepts through prompts or system instructions, but no rigorous benchmark exists for evaluating them on visual tasks. We introduce VLM-UnBench, the first benchmark for training-free visual concept unlearning in VLMs. It covers four forgetting levels, 7 source datasets, and 11 concept axes, and pairs a three-level probe taxonomy with five evaluation conditions to separate genuine forgetting from instruction compliance. Across 8 evaluation settings and 13 VLM configurations, realistic unlearning prompts leave forget accuracy near the no-instruction baseline; meaningful reductions appear only under oracle conditions that disclose the target concept to the model. Object and scene concepts are the most resistant to suppression, and stronger instruction-tuned models remain capable despite explicit forget instructions. These results expose a clear gap between prompt-level suppression and true visual concept erasure.

视觉模型概念遗忘无训练方法评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。