arXiv:2608.01849cs.AI2026-08

提出新方法修复遗忘模型的隐性知识漏洞,提升安全性和性能。

Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models

论文配图:Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models
图 1 · 摘自论文原文
  • 通过锚定激活过滤和抽象增强,选择性保护通用模式。
  • 在SafeEraser上恢复98%原始响应质量,攻击成功率0%。
  • 揭示现有评估盲区,适合关注模型安全与可靠性的研究者。

机器遗忘为消除多模态大语言模型(MLLM)中的不安全内容提供了有前景的途径,但确保遗忘精度仍是持续挑战。当前的MLLM遗忘评估范式存在关键盲点:通过与遗忘集表示相距较远的基准评估模型效用,无法捕捉知识漏洞——即对与遗忘集具有通用模式的良性邻近输入的严重退化。为探测未遗忘MLLM中的知识漏洞,我们构建了一个基准,捕获与遗忘集共享通用模式的良性输入上的意外退化,并通过受控实验确认其是常用方法的系统性后果。为进一步弥合这一差距,我们提出选择性保护锚定正则化(SPAR),通过锚定激活过滤保护通用模式,同时通过实体抽象增强强化它们。在SafeEraser上的实验表明,SPAR相较于标准基线恢复超过98%的原始响应质量(基线低于50%),同时实现0.00%攻击成功率且具备竞争力的模型效用。这些结果凸显了更细粒度评估对于可信MLLM遗忘的必要性。

原文摘要 · Abstract (English)

Machine unlearning offers a promising approach to remove unsafe content from Multimodal Large Language Models (MLLMs), yet ensuring the precision of unlearning remains a persistent challenge. One reason is that current MLLM unlearning evaluation paradigms suffer from a critical blind spot: they assess model utility through benchmarks whose representations are distant from the forget set, failing to capture knowledge holes---severe degradation on benign adjacent inputs. To probe knowledge holes in unlearned MLLMs, we construct a benchmark that captures unintended degradation on benign inputs sharing generic patterns with the forget set, and confirm through controlled experiments that they are a systematic consequence of commonly used approaches. Furthermore, to bridge this gap, we propose Selective Protection with Anchored Regularization, which protects generic patterns via anchored activation filtering while reinforcing them through entity-abstracted enhancement. Our experiments on SafeEraser demonstrate that SPAR recovers over 98% of vanilla response quality compared to below 50% for standard baselines---while achieving 0.00% attack success rate and competitive model utility. These results underscore the necessity of more fine-grained evaluation for trustworthy MLLM unlearning.

模型遗忘多模态安全性评估改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。