新基准可检测视觉模型推理是否真实依赖图像细节。
Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images
- 设计可验证过程的基准ViEBench,分感知与推理难度
- 200张高清图带专家标注证据,支持细粒度评估
- 发现模型常答对但用错区域,或找对却不会用
尽管视觉语言模型(VLMs)在“图像思维”能力上取得显著进展,但准确评估其推理过程的真实性仍是关键挑战。现有基准主要依赖结果准确性,无法判断模型是否真正利用细粒度视觉线索进行多步推理。为此,我们提出ViEBench,一个可验证过程的基准,用于评估忠实的视觉推理。该基准包含200张多场景高分辨率图像,配有专家标注的视觉证据,按感知与推理难度分类,其中推理任务需结合局部视觉细节与先验知识。为建立全面评估标准,我们引入双轴矩阵,通过四个诊断象限提供细粒度指标,实现对不同复杂度任务下模型行为的透明诊断。实验发现:(1) 模型有时能得出正确答案,但依据的是无关区域;(2) 模型可能准确定位正确证据,却未能有效利用以得出准确结论。结果表明,ViEBench可作为更可解释、实用的基准,全面评估代理型VLM的有效性。代码将发布于:https://github.com/Xuchen-Li/ViEBench。
原文摘要 · Abstract (English)
Despite the remarkable progress of Vision-Language Models (VLMs) in adopting "Thinking-with-Images" capabilities, accurately evaluating the authenticity of their reasoning process remains a critical challenge. Existing benchmarks mainly rely on outcome-oriented accuracy, lacking the capability to assess whether models can accurately leverage fine-grained visual cues for multi-step reasoning. To address these limitations, we propose ViEBench, a process-verifiable benchmark designed to evaluate faithful visual reasoning. Comprising 200 multi-scenario high-resolution images with expert-annotated visual evidence, ViEBench uniquely categorizes tasks by difficulty into perception and reasoning dimensions, where reasoning tasks require utilizing localized visual details with prior knowledge. To establish comprehensive evaluation criteria, we introduce a dual-axis matrix that provides fine-grained metrics through four diagnostic quadrants, enabling transparent diagnosis of model behavior across varying task complexities. Our experiments yield several interesting observations: (1) VLMs can sometimes produce correct final answers despite grounding on irrelevant regions, and (2) they may successfully locate the correct evidence but still fail to utilize it to reach accurate conclusions. Our findings demonstrate that ViEBench can serve as a more explainable and practical benchmark for comprehensively evaluating the effectiveness agentic VLMs. The codes will be released at: https://github.com/Xuchen-Li/ViEBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。