测试顶级视觉语言模型的跨区域视觉推理能力,发现其表现远低于人类。
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
- 设计三类非局部视觉推理任务:对比感知、扫视搜索和连续追踪。
- 主流模型在部分任务上准确率仅略高于随机水平,人类可轻松完成。
- 揭示当前模型虽有视觉识别能力,但缺乏核心视觉推理机制,适合研究者参考。
视觉语言模型(VLMs)在复杂视觉任务如视觉问答(VQA)和图表理解中表现优异,但近期研究表明它们在基础感知测试中存在短板。本文评估了领先VLMs在非局部视觉推理方面的能力——即需要整合图像中多个、可能相距较远区域证据的推理。我们隔离出三种非局部视觉形式:比较感知(需在工作记忆中保持并比较两幅图像)、扫视搜索(需离散、基于证据地跳转定位目标)以及连续视觉搜索(需沿连续轮廓追踪)。旗舰模型(如GPT-5、Gemini 2.5 Pro、Claude Sonnet 4)即使在先前的初级视觉基准上表现良好,仍在此类任务中失败,在两个对人类而言极为简单的任务变体上准确率仅略高于随机水平。我们的结构化评测套件可用于检验VLM是否具备类似人类的视觉算法能力。结果表明,尽管视觉敏锐度有所提升,当前模型仍缺乏核心视觉推理能力。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) excel at complex visual tasks such as VQA and chart understanding, yet recent work suggests they struggle with simple perceptual tests. We present an evaluation of vision-language models' capacity for nonlocal visual reasoning: reasoning that requires chaining evidence collected from multiple, possibly distant regions of an image. We isolate three distinct forms of nonlocal vision: comparative perception, which demands holding two images in working memory and comparing them; saccadic search, which requires making discrete, evidence-driven jumps to locate successive targets; and smooth visual search, which involves following a continuous contour. Flagship models (e.g., GPT-5, Gemini 2.5 Pro, Claude Sonnet 4), even those that perform well on prior primitive-vision benchmarks, fail these tests and barely exceed random accuracy on two variants of our tasks that are trivial for humans. Our structured evaluation suite allows us to test whether VLMs can perform visual algorithms similar to those used by humans. Our findings show that despite gains in raw visual acuity, current models lack core visual reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。