arXiv:2606.14795cs.CV2026-06中稿 · ICML

现有视觉模型缺乏主动探索能力,难以发现隐藏视觉线索。

Position: The Systemic Lack of Agency in Visual Reasoning

论文配图:Position: The Systemic Lack of Agency in Visual Reasoning
图 1 · 摘自论文原文
  • 提出主动视觉推理新评测基准V-IRD,要求模型自主探索图像
  • 主流模型在隐含线索利用上表现差,即使识别能力强也难发现关键信息
  • 适合关注模型真实推理能力、非检索式视觉理解的研究者

本文指出,当前视觉语言模型(VLMs)存在系统性代理缺失,限制了其隐含推理能力。隐含推理指模型能自主发现并利用隐藏视觉证据以填补信息空白,而非仅依赖显式提示。这种能力是人类视觉理解的基础。我们发现,现有模型多将视觉推理视为被动语义检索,而非依赖自主视觉探索的主动情境推理,导致多数基准仅评估被动能力,忽略主动探索维度。为此,我们提出视觉隐含推理诊断基准(V-IRD),要求模型完全通过自主视觉分析得出答案。结果表明,尽管顶尖模型具备强大识别能力,却仍难以有效利用参考对象或关注需自导探索的视觉证据。这揭示了强语义识别与主动视觉探索之间的关键鸿沟。

原文摘要 · Abstract (English)

This paper argues that a systemic lack of Agency constrains the implicit reasoning capabilities of current Vision-Language Models (VLMs). Implicit reasoning refers to the ability to autonomously discover and utilize hidden visual evidence to bridge information gaps, rather than merely relying on explicitly specified targets. This capacity underlies human visual understanding and everyday reasoning. We argue that this limitation arises from a tendency to approach visual reasoning primarily as passive semantic retrieval, rather than as active, situated reasoning that depends on autonomous visual exploration. As a result, most existing benchmarks primarily assess Passive Capacity, leaving this aspect of reasoning largely unmeasured. To address this gap, we introduce the Visual Implicit Reasoning Diagnosing Benchmark (V-IRD), which targets this missing quadrant by requiring models to derive answers strictly through autonomous visual analysis. Our results show that, despite strong retrieval abilities, prominent VLMs struggle to utilize reference objects and to attend to visual evidence that requires self-directed inquiry. Simply put, strong semantic recognition does not equate to active visual exploration, revealing a critical gap in current VLMs. More information can be found at https://haoychen.github.io/Implicit-Reasoning/

视觉推理主动探索模型评估隐含线索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。