arXiv:2608.03791cs.AI2026-08

提出首个跨模态遗忘评估基准,揭示文本遗忘难传到视觉模态。

Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation

论文配图:Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation
图 1 · 摘自论文原文
  • 构建真实实体关联图文与知识图谱的跨模态遗忘测试集
  • 发现多模态遗忘在文本上有效,但难传到视觉和跨模态场景
  • 强调必须用跨模态评估才能真实衡量模型遗忘效果

视觉-语言模型(VLMs)可能记忆敏感、版权或有害知识,移除这些知识对构建可信AI至关重要。现有研究多集中于单模态遗忘,跨模态一致性仍不足。为此,我们提出UNLINK-VL,一个真实世界下VLM跨模态知识遗忘的评估基准。在原始遗忘与保留数据不可用的后处理遗忘设置中,选取可视觉识别的真实实体作为遗忘目标,关联对应图像及来自Wikidata的一跳和多跳事实。基准包含四个互补子集:直接遗忘效果、关系知识传播、非目标相关知识保留、对语义等价查询的鲁棒性。在纯文本和多模态遗忘设置下训练模型,并在文本、视觉和跨模态场景中评估遗忘效果与保留能力。实验表明跨模态转移存在显著不对称:多模态遗忘在文本评估中有效,但文本仅遗忘在视觉和跨模态场景中表现差。同时,各方法基本保持模型通用能力。结果表明,仅依赖单模态(尤其是文本)评估会严重高估遗忘效果,亟需跨模态遗忘与评估。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpora. Removing such knowledge is essential for building trustworthy AI systems. However, existing studies primarily focus on forgetting within individual modalities. Although recent work has begun to explore cross-modal consistency in unlearning, the cross-modal transfer of real-world knowledge unlearning remains insufficiently studied. To address this gap, we introduce UNLINK-VL, a real-world benchmark for cross-modal knowledge unlearning in VLMs. Under a post-hoc unlearning setting in which the original forget and retain corpora are unavailable, UNLINK-VL selects visually identifiable real-world entities as unlearning targets and associates them with corresponding images and one-hop and multi-hop facts derived from Wikidata. The benchmark comprises four complementary subsets that evaluate direct forgetting of target knowledge, the propagation of forgetting through relational knowledge, the preservation of related non-target knowledge, and robustness to semantically equivalent queries. We train models under text-only and multimodal unlearning settings and evaluate forgetting effectiveness and retained utility across textual, visual, and cross-modal scenarios. Extensive experiments reveal a pronounced asymmetry in cross-modal transfer: multimodal unlearning remains effective under textual evaluation, whereas text-only unlearning transfers poorly to visual and cross-modal scenarios. Meanwhile, the evaluated methods largely preserve the models' general capabilities. These findings demonstrate that relying solely on intra-modal evaluation, particularly text-only evaluation, may substantially overestimate the effectiveness of knowledge unlearning in VLMs, underscoring the need for cross-modal unlearning and evaluation.

跨模态知识遗忘评估基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。