arXiv:2510.17771cs.AIcs.CV2025-10被引 37

发现视觉模型'看得到却不信',干预后准确率显著提升

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

  • 通过分析注意力层动态,发现深层网络能定位正确视觉区域
  • 多数错误答案仍基于正确视觉信息,存在'见而未信'现象
  • 无需训练的注意力掩码干预,跨模型有效提升准确率

视觉语言模型在多模态任务中表现强劲,但仍会在视觉证据存在时出错。本文系统探究这些失败是否源于未感知证据或未能有效利用。通过分析逐层注意力动态,发现浅层主要关注文本,深层则稀疏但可靠地聚焦于局部证据区域。令人意外的是,模型在输出错误答案时,往往已正确感知视觉证据,这种现象被称为“看见但不相信”,广泛存在于主流视觉语言模型家族中。基于此,我们提出一种推理时干预方法:通过选择性注意力掩码突出深层证据区域,无需训练即可一致提升多个模型(包括LLaVA、Qwen、Gemma和InternVL)的准确率。结果表明,模型内部已编码可靠证据但使用不足,显式强化此类信号可弥合感知与推理间的差距,推动对视觉语言模型诊断与可靠性理解的进展。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) achieve strong results on multimodal tasks such as visual question answering, yet they can still fail even when the correct visual evidence is present. In this work, we systematically investigate whether these failures arise from not perceiving the evidence or from not leveraging it effectively. By examining layer-wise attention dynamics, we find that shallow layers focus primarily on text, while deeper layers sparsely but reliably attend to localized evidence regions. Surprisingly, VLMs often perceive the visual evidence when outputting incorrect answers, a phenomenon we term ``seeing but not believing'' that widely exists in major VLM families. Building on this, we introduce an inference-time intervention that highlights deep-layer evidence regions through selective attention-based masking. It requires no training and consistently improves accuracy across multiple families, including LLaVA, Qwen, Gemma, and InternVL. These results show that VLMs encode reliable evidence internally but under-utilize it, making such signals explicit can bridge the gap between perception and reasoning, advancing the diagnostic understanding and reliability of VLMs.

视觉语言模型注意力机制模型可靠性推理干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。