arXiv:2502.20808cs.AI2025-02CVPR被引 56

构建多视觉数学推理数据集,评估大模型在真实场景下的解题能力

MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts

  • 设计2009道含多图与文本的数学题,覆盖K-12真实场景
  • 模型在多视觉任务中表现远低于人类,差距显著
  • 揭示模型在多图协同理解中的系统性错误模式

多模态大语言模型(MLLMs)在单视觉场景下的数学推理已展现潜力,但现有基准大多局限于单一图像,与真实世界中多视觉数学应用不符。为此,我们提出MV-MATH:一个精心构建的高质量数学问题数据集,包含2,009道题目。每道题融合多张图像与文本,源自真实的K-12教育场景,并配有详尽标注。数据集涵盖选择题、自由作答和多步推理题,覆盖11个学科领域及3个难度等级,可全面评估MLLMs在多视觉环境中的数学推理能力。实验表明,当前MLLMs在多视觉数学任务中面临严峻挑战,在MV-MATH上的表现与人类存在明显差距。我们进一步分析了不同模型的表现与错误模式,深入揭示其在多视觉语境下的推理局限。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown promising capabilities in mathematical reasoning within visual contexts across various datasets. However, most existing multimodal math benchmarks are limited to single-visual contexts, which diverges from the multi-visual scenarios commonly encountered in real-world mathematical applications. To address this gap, we introduce MV-MATH: a meticulously curated dataset of 2,009 high-quality mathematical problems. Each problem integrates multiple images interleaved with text, derived from authentic K-12 scenarios, and enriched with detailed annotations. MV-MATH includes multiple-choice, free-form, and multi-step questions, covering 11 subject areas across 3 difficulty levels, and serves as a comprehensive and rigorous benchmark for assessing MLLMs' mathematical reasoning in multi-visual contexts. Through extensive experimentation, we observe that MLLMs encounter substantial challenges in multi-visual math tasks, with a considerable performance gap relative to human capabilities on MV-MATH. Furthermore, we analyze the performance and error patterns of various models, providing insights into MLLMs' mathematical reasoning capabilities within multi-visual settings.

多模态数学推理视觉理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。