构建视觉推理评测基准,揭示多模态大模型真实视觉理解短板
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- 设计1000个经人工验证的跨类别视觉推理题,涵盖数量变化、空间关系等六类
- 主流模型平均准确率低于30%,远低于人类51.4%水平,仅略高于随机基线25%
- 提供训练数据集与强化学习基线,助力提升模型视觉推理能力
视觉推理是人类智能的核心组成部分,也是先进多模态模型的关键能力。然而,当前多模态大语言模型(MLLMs)的推理评估大多依赖文本描述,允许语言层面的捷径,无法真正衡量以视觉为中心的推理能力。为此,我们提出VisuLogic:一个包含1000个经人工验证问题的基准,覆盖六类任务(如数量变化、空间关系、属性比较等),可从多角度评估MLLMs的视觉推理能力。我们在该基准上评估了主流的MLLMs并分析其结果,发现大多数模型准确率低于30%,仅略高于25%的随机基线,远低于人类51.4%的表现,暴露出显著的能力差距。此外,我们还提供了补充训练数据集和强化学习基线,以支持后续研究进展。
原文摘要 · Abstract (English)
Visual reasoning is a core component of human intelligence and a critical capability for advanced multimodal models. Yet current reasoning evaluations of multimodal large language models (MLLMs) often rely on text descriptions and allow language-based reasoning shortcuts, failing to measure genuine vision-centric reasoning. To address this, we introduce VisuLogic: a benchmark of 1,000 human-verified problems across six categories (e.g., quantitative shifts, spatial relations, attribute comparisons). These various types of questions can be evaluated to assess the visual reasoning capabilities of MLLMs from multiple perspectives. We evaluate leading MLLMs on this benchmark and analyze their results to identify common failure modes. Most models score below 30% accuracy-only slightly above the 25% random baseline and far below the 51.4% achieved by humans-revealing significant gaps in visual reasoning. Furthermore, we provide a supplementary training dataset and a reinforcement-learning baseline to support further progress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。