arXiv:2508.04852cs.CV2025-08被引 13

评测大模型对细微视觉线索的推理能力,发现现有模型严重不足。

VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence

  • 设计细粒度视觉线索检测与推理框架,聚焦仅占0.25%图像面积的微小线索
  • 在6类复杂推理任务中,模型提取细微线索能力普遍弱于人类水平
  • 适合关注视觉理解深度、多模态推理的研究者和开发者

随着多模态大模型(MLLMs)快速发展,评估其视觉能力变得愈发重要。现有基准主要分为两类:基础感知类(如“图中有什么?”),侧重局部细节但缺乏深层推理;主流推理类则聚焦显著图像元素,难以评估需精细分析的细微线索。然而,真正深入的视觉理解依赖于对隐蔽、微小局部细节的解析,这些细节虽仅占图像平均0.25%面积,却常蕴含关键信息。为此,我们提出VER-Bench,一个全新评估框架,用于衡量MLLMs在识别细粒度视觉线索(平均占0.25%图像面积)并结合常识进行复杂推理的能力。该框架包含374个精心设计的问题,覆盖地理、时间、情境、意图、系统状态与符号推理六类,每个问题均配有结构化证据:视觉线索及其推导出的推理依据。实验揭示当前模型在提取细微视觉证据和构建证据链方面存在明显缺陷,凸显了提升模型细粒度视觉理解与推理能力的必要性。数据集及相关材料已公开于https://github.com/verbta/ACMMM-25-Materials。

原文摘要 · Abstract (English)

With the rapid development of MLLMs, evaluating their visual capabilities has become increasingly crucial. Current benchmarks primarily fall into two main types: basic perception benchmarks, which focus on local details but lack deep reasoning (e.g., "what is in the image?"), and mainstream reasoning benchmarks, which concentrate on prominent image elements but may fail to assess subtle clues requiring intricate analysis. However, profound visual understanding and complex reasoning depend more on interpreting subtle, inconspicuous local details than on perceiving salient, macro-level objects. These details, though occupying minimal image area, often contain richer, more critical information for robust analysis. To bridge this gap, we introduce the VER-Bench, a novel framework to evaluate MLLMs' ability to: 1) identify fine-grained visual clues, often occupying on average just 0.25% of the image area; 2) integrate these clues with world knowledge for complex reasoning. Comprising 374 carefully designed questions across Geospatial, Temporal, Situational, Intent, System State, and Symbolic reasoning, each question in VER-Bench is accompanied by structured evidence: visual clues and question-related reasoning derived from them. VER-Bench reveals current models' limitations in extracting subtle visual evidence and constructing evidence-based arguments, highlighting the need to enhance models's capabilities in fine-grained visual evidence extraction, integration, and reasoning for genuine visual understanding and human-like analysis. Dataset and additional materials are available https://github.com/verbta/ACMMM-25-Materials.

多模态视觉推理细粒度分析评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。