arXiv:2605.22903cs.CVcs.AI2026-05中稿 · GRAIL-V: Grounded …被引 3

现有视觉语言模型的准确率可能虚高,因它们依赖文本线索而非真实视觉细节。

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

论文配图:Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?
图 1 · 摘自论文原文
  • 通过移除图像关键区域,测试模型对视觉信息的依赖程度。
  • 即使图像信息丢失超过一半,模型准确率仍基本不变。
  • 适合关注模型真实理解能力、而非表面表现的研究者阅读。

基准测试准确率常被默认反映视觉语言模型(VLMs)对视觉内容的扎根理解,但其是否真正依赖视觉证据尚不明确。我们观察到,在一个广泛使用的幻觉检测基准上,大幅移除图像标记后,模型性能仅轻微下降,由此系统性地研究了多款开源VLMs中的这一差异。分析涵盖全局视觉退化、局部遮挡、问题重述、答案空间扩展及超越标准准确率的决策层面分析,并结合层间视觉标记几何结构分析。结果表明,尽管VLMs确实使用视觉输入,但其预测对细粒度视觉证据的敏感性远低于标准准确率所暗示的程度。即使最终输出未变,内部支持正确答案的信号可能已减弱。进一步的表示层分析显示,深层中视觉标记间的相似性逐渐增加,可能解释该现象。综合来看,当前基准无法可靠评估VLMs的细粒度视觉定位能力。

原文摘要 · Abstract (English)

Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly reflect reliance on visual evidence. Motivated by a surprising observation that removing a substantial fraction of image tokens only degrades model performance very slightly on a widely used hallucination benchmark, we systematically investigate this mismatch in a set of open-source VLMs. Our analysis spans multiple levels of granularity, spanning global visual degradation, localized occlusion, question reformulation, answer-space expansion, and decision-level analyses beyond standard accuracy. We further complement these behavioral results with a layer-wise analysis of vision-token geometry. Throughout the experiments, we find that although VLMs do incorporate visual input, their predictions are less sensitive to the loss of fine-grained visual evidence that standard accuracy should have suggested. Even when the final prediction remains unchanged, the model's internal support for the correct answer may already be weakened. We further complement a representation-level analysis, which shows increasing similarity among visual tokens in deeper layers, providing a possible explanation for our findings. Together, these results suggest that current benchmarks are not sufficient to reliably evaluate fine-grained visual grounding in VLMs.

视觉语言模型基准测试视觉理解可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。