arXiv:2604.06250cs.CVcs.AI2026-04

诊断视觉模型'看懂却想不对'的根源,揭示推理短板。

DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs

  • 设计五种输入模式,拆解视觉理解与推理能力
  • 发现开源模型在自述图像后推理更差,暴露整合瓶颈
  • 闭源模型无此差距,反映其整合能力领先

当被要求描述分子图时,视觉语言模型能正确识别‘苯环带羟基’;但若需推理,则回答错误。这表明模型能‘看见’却无法‘思考’所见内容。我们称此为感知-整合鸿沟:视觉信息虽成功提取,但在下游推理中丢失,而传统单配置基准将感知与整合混为一谈。为此,我们提出DISSECT,一个包含12,000个问题的诊断基准,涵盖化学(7,000题)与生物(5,000题)。每个问题在五种输入模式下评估:视觉+文本、纯文本、纯视觉、人类先知和一种新型模型先知——模型先用自己的语言描述图像,再基于描述推理。该设计可分解性能至语言先验利用、视觉提取、感知保真度与整合有效性。评估18个VLM后发现:(1) 化学任务的语言先验可利用性显著低于生物,证实分子视觉内容是更严格的视觉推理考验;(2) 开源模型在使用自身描述进行推理时表现优于直接使用图像,暴露系统性整合瓶颈;(3) 闭源模型无此差距,说明感知与整合的融合是开源自闭源多模态能力的核心分水岭。模型先知协议对模型与基准均无依赖,可事后应用于任意VLM评估以诊断整合失败。

原文摘要 · Abstract (English)

When asked to describe a molecular diagram, a Vision-Language Model correctly identifies ``a benzene ring with an -OH group.'' When asked to reason about the same image, it answers incorrectly. The model can see but it cannot think about what it sees. We term this the perception-integration gap: a failure where visual information is successfully extracted but lost during downstream reasoning, invisible to single-configuration benchmarks that conflate perception with integration under one accuracy number. To systematically expose such failures, we introduce DISSECT, a 12,000-question diagnostic benchmark spanning Chemistry (7,000) and Biology (5,000). Every question is evaluated under five input modes -- Vision+Text, Text-Only, Vision-Only, Human Oracle, and a novel Model Oracle in which the VLM first verbalizes the image and then reasons from its own description -- yielding diagnostic gaps that decompose performance into language-prior exploitation, visual extraction, perception fidelity, and integration effectiveness. Evaluating 18~VLMs, we find that: (1) Chemistry exhibits substantially lower language-prior exploitability than Biology, confirming molecular visual content as a harder test of genuine visual reasoning; (2) Open-source models consistently score higher when reasoning from their own verbalized descriptions than from raw images, exposing a systematic integration bottleneck; and (3) Closed-source models show no such gap, indicating that bridging perception and integration is the frontier separating open-source from closed-source multimodal capability. The Model Oracle protocol is both model and benchmark agnostic, applicable post-hoc to any VLM evaluation to diagnose integration failures.

多模态视觉推理诊断评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。