arXiv:2503.11557cs.CV2025-03被引 25

评测大模型视觉推理能力,发现现有模型依赖语言而非真实看图思考。

VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity

  • 设计新基准VERIFY,让模型仅凭图像推理,减少文字干扰。
  • 发现顶尖模型在复杂推理任务中准确率不足50%,暴露严重缺陷。
  • 提供人类标注的推理路径,适合研究模型决策过程的学者使用。

视觉推理是人类认知的核心,使个体能够解读并抽象理解环境。尽管近期多模态大语言模型(MLLMs)在语言和视觉-语言任务上表现出色,但现有基准主要衡量识别能力,未能充分评估真正的视觉推理水平。为弥补这一关键差距,我们提出VERIFY,一个专门设计用于隔离并严格评估前沿MLLM视觉推理能力的基准。该基准要求模型主要基于视觉信息进行推理,同时提供极少量文本上下文,以减少对领域知识和语言偏见的依赖。每个问题均配有由人工标注的推理路径,是首个深入评估模型决策过程的基准。此外,我们提出新型度量方法,超越单纯准确率,揭示当前模型推理模式中的显著失衡。对主流MLLM的全面测评揭示了显著局限性,强调需采用平衡且全面的方法来提升感知与推理能力。更多试用和展示请访问项目页面(https://verify-eqh.pages.dev/)。

原文摘要 · Abstract (English)

Visual reasoning is central to human cognition, enabling individuals to interpret and abstractly understand their environment. Although recent Multimodal Large Language Models (MLLMs) have demonstrated impressive performance across language and vision-language tasks, existing benchmarks primarily measure recognition-based skills and inadequately assess true visual reasoning capabilities. To bridge this critical gap, we introduce VERIFY, a benchmark explicitly designed to isolate and rigorously evaluate the visual reasoning capabilities of state-of-the-art MLLMs. VERIFY compels models to reason primarily from visual information, providing minimal textual context to reduce reliance on domain-specific knowledge and linguistic biases. Each problem is accompanied by a human-annotated reasoning path, making it the first to provide in-depth evaluation of model decision-making processes. Additionally, we propose novel metrics that assess visual reasoning fidelity beyond mere accuracy, highlighting critical imbalances in current model reasoning patterns. Our comprehensive benchmarking of leading MLLMs uncovers significant limitations, underscoring the need for a balanced and holistic approach to both perception and reasoning. For more teaser and testing, visit our project page (https://verify-eqh.pages.dev/).

多模态视觉推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。