arXiv:2511.13243cs.LGcs.AI2025-11AAAI

发现多模态编辑中视觉盲区问题并提出缓解方法

Uncovering and Mitigating Transient Blindness in Multimodal Model Editing

  • 构建三维度局部性评估框架,覆盖随机、无图、一致图像场景
  • 发现编辑后模型短暂忽视视觉信息,文本过度拟合现象
  • 设计自适应对抗损失,提升跨模态平衡,平均改善17%

多模态模型编辑(MMED)旨在修正多模态模型中的错误知识。现有评估方法沿用文本模型编辑的范式,依赖低相似度或随机输入,导致评估结果虚高,掩盖了过拟合问题。本文提出一个全面的局部性评估框架,涵盖随机图像局部性、无图局部性与一致图像局部性三个关键维度,通过七种不同数据类型实现对多模态编辑的细致结构化分析。我们引入De-VQA动态视觉问答评估,揭示了一种称为瞬时盲区(transient blindness)的现象:模型在编辑后过度拟合与编辑内容相似的文本,却忽略视觉信息。词元分析显示,编辑对文本词元的影响远高于视觉特征。为此,我们提出局部性感知对抗损失,以平衡跨模态表示。实证结果表明,该方法持续优于现有基线,在降低瞬时盲区和提升局部性方面平均提升17%。

原文摘要 · Abstract (English)

Multimodal Model Editing (MMED) aims to correct erroneous knowledge in multimodal models. Existing evaluation methods, adapted from textual model editing, overstate success by relying on low-similarity or random inputs, obscure overfitting. We propose a comprehensive locality evaluation framework, covering three key dimensions: random-image locality, no-image locality, and consistent-image locality, operationalized through seven distinct data types, enabling a detailed and structured analysis of multimodal edits. We introduce De-VQA, a dynamic evaluation for visual question answering, uncovering a phenomenon we term transient blindness, overfitting to edit-similar text while ignoring visuals. Token analysis shows edits disproportionately affect textual tokens. We propose locality-aware adversarial losses to balance cross-modal representations. Empirical results demonstrate that our approach consistently outperforms existing baselines, reducing transient blindness and improving locality by 17% on average.

多模态编辑视觉盲区跨模态对齐模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。