arXiv:2608.06769cs.CV2026-08

构建首个图文大模型图结构推理综合评测基准

GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models

论文配图:GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
图 1 · 摘自论文原文
  • 设计图中心图像编辑策略,动态测试模型视觉推理能力
  • 提出过程敏感评分指标,超越传统答案准确率
  • 支持单图与配对图像场景,覆盖真实复杂推理任务

当前多模态大模型在跨视觉语言任务中进展显著,亟需更复杂的评估基准。现有评测难以检验模型对结构化视觉信息的真实推理能力。视觉图推理(VGR)为该挑战提供理想测试平台,要求模型融合感知、结构理解与多步推理。然而,先前的VGR基准常将任务简化为视觉感知后接文本推理,仅限单图设置,依赖答案仅指标,且缺乏真实图结构场景。为此,我们提出GraphVerse,一个统一基准,联合评估多模态大模型在单图与配对图像场景下的感知、视觉推理与文本图推理能力。其核心是一套图中心图像编辑(GIE)策略,可在保持语义不变的前提下修改图图像,使其成为主动的视觉推理测试。我们进一步提出VGR-Score,一种过程敏感指标,可评估推理质量而不仅看最终答案。大量实验揭示当前多模态大模型在图推理中的多项局限,同时验证了GIE策略的有效性及GraphVerse向更广泛多模态推理能力的迁移潜力。代码已开源。

原文摘要 · Abstract (English)

Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more challenging benchmarks. Yet existing evaluations still provide limited insight into whether these models can truly reason over structured visual information. Visual Graph Reasoning (VGR) offers a compelling testbed for this challenge, requiring models to integrate perception, structural understanding, and multi-step reasoning over graph-based visual inputs. However, prior VGR benchmarks often reduce the task to visual perception followed by text-based reasoning, restrict evaluation to single-image settings, rely on answer-only metrics, and underrepresent realistic graph-centric scenarios. To bridge the gap, we introduce GraphVerse, a unified benchmark that jointly evaluates perception, visual reasoning, and text-based graph reasoning in MLLMs under both single-image and paired-image settings. At its core is a suite of Graph-centric Image Editing (GIE) strategies that modify graph images while preserving their semantics, turning them into active tests of visual reasoning. We further propose VGR-Score, a process-sensitive metric that evaluates reasoning quality beyond final-answer accuracy. Extensive experiments reveal several key limitations of current MLLMs in VGR, while also validating the effectiveness of GIE strategies and the transferability of GraphVerse to broader multimodal reasoning capabilities. The code is available at https://github.com/sunyuanfu/GraphVerse.

图推理多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。