arXiv:2505.10541cs.CV2025-05被引 2

发现多模态模型答对题却没看懂图,用注意力分析揭示隐性误解

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis

  • 通过拆解注意力机制,发现模型越来越聚焦正确答案对应的图像
  • 提出注意力准确率新指标,能无偏评估视觉理解能力
  • 适用于多模态和单模态场景,可检测模型是否真看懂了图

近年来,多模态大语言模型(MLLMs)在理解多图信息方面取得进展。然而,现有评测主要关注答案正确性,忽略了模型是否真正理解视觉输入。为此,我们定义了隐性视觉误解(IVM),即模型给出正确答案但未充分理解视觉内容。通过分析因果注意力模块,我们发现随着网络层加深,注意力分布逐渐集中到与正确答案相关的图像上。基于此,提出一种与尺度无关的指标——注意力准确率(attention accuracy),并构建新基准以量化IVM。该指标通过内部机制直接评估视觉理解,对位置偏差具有鲁棒性,提升评估可靠性。进一步扩展至细粒度分析,在单模态场景中也验证了其有效性,证明方法具备广泛适用性与通用性。

原文摘要 · Abstract (English)

Recent advancements have enhanced the capability of Multimodal Large Language Models (MLLMs) to comprehend multi-image information. However, existing benchmarks primarily evaluate answer correctness, overlooking whether models genuinely comprehend the visual input. To address this, we define implicit visual misunderstanding (IVM), where MLLMs provide correct answers without fully comprehending the visual input. Through our analysis, we decouple the visual and textual modalities within the causal attention module, revealing that attention distribution increasingly converges on the image associated with the correct answer as the network layers deepen. This insight leads to the introduction of a scale-agnostic metric, \textit{attention accuracy}, and a novel benchmark for quantifying IVMs. Attention accuracy directly evaluates the model's visual understanding via internal mechanisms, remaining robust to positional biases for more reliable assessments. Furthermore, we extend our approach to finer granularities and demonstrate its effectiveness in unimodal scenarios, underscoring its versatility and generalizability.

多模态注意力分析模型理解评测方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。