arXiv:2504.18589cs.CV2025-04被引 9

构建多图依赖数学推理基准,揭示视觉语言模型在基础数学上的短板

Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency

  • 设计需跨多图整合信息的数学题,测试视觉与数学推理融合能力
  • 26个主流模型平均准确率不足50%,顶尖模型也仅达49.8%
  • 适合关注多模态推理、AGI进展的研究者与开发者

大型视觉语言模型(LVLMs)在物体识别、图像描述和视觉问答等任务中已接近人类水平,但现有评测多聚焦领域知识,忽视对基础数学元素与视觉概念的推理能力评估。本文指出,当前基准缺乏对显式视觉依赖的数学问题测评,这类问题要求模型从多张图中提取、整合并推理信息,结合常识知识,是迈向通用人工智能的关键。为此,我们提出VCBENCH,一个涵盖六类认知领域、共1,720道题目的多模态数学推理基准,包含6,697张图像(平均每题3.9张),支持多图推理。我们在该基准上评估26个先进LVLMs,发现性能差距显著,即使最优模型准确率也未超过50%。结果表明,视觉-数学融合仍面临巨大挑战,为未来模型改进提供方向。项目地址:https://alibaba-damo-academy.github.io/VCBench/

原文摘要 · Abstract (English)

Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, captioning, and visual question answering. However, current benchmarks typically focus on knowledge-centric evaluations that assess domain-specific expertise, often neglecting the core ability to reason about fundamental mathematical elements and visual concepts. We identify a gap in evaluating elementary-level math problems, which rely on explicit visual dependencies-requiring models to discern, integrate, and reason across multiple images while incorporating commonsense knowledge, all of which are crucial for advancing toward broader AGI capabilities. To address this gap, we introduce VCBENCH, a comprehensive benchmark for multimodal mathematical reasoning with explicit visual dependencies. VCBENCH includes 1,720 problems across six cognitive domains, featuring 6,697 images (averaging 3.9 per question) to ensure multi-image reasoning. We evaluate 26 state-of-the-art LVLMs on VCBENCH, revealing substantial performance disparities, with even the top models unable to exceed 50% accuracy. Our findings highlight the ongoing challenges in visual-mathematical integration and suggest avenues for future LVLM advancements. The project can be found at https://alibaba-damo-academy.github.io/VCBench/.

多模态推理数学推理视觉语言模型基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。