构建多视觉数学推理数据集,评估大模型在真实场景下的解题能力
MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts
- 设计2009道含多图与文本的数学题,覆盖K-12真实场景
- 模型在多视觉任务中表现远低于人类,差距显著
- 揭示模型在多图协同理解中的系统性错误模式
多模态大语言模型(MLLMs)在单视觉场景下的数学推理已展现潜力,但现有基准大多局限于单一图像,与真实世界中多视觉数学应用不符。为此,我们提出MV-MATH:一个精心构建的高质量数学问题数据集,包含2,009道题目。每道题融合多张图像与文本,源自真实的K-12教育场景,并配有详尽标注。数据集涵盖选择题、自由作答和多步推理题,覆盖11个学科领域及3个难度等级,可全面评估MLLMs在多视觉环境中的数学推理能力。实验表明,当前MLLMs在多视觉数学任务中面临严峻挑战,在MV-MATH上的表现与人类存在明显差距。我们进一步分析了不同模型的表现与错误模式,深入揭示其在多视觉语境下的推理局限。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have shown promising capabilities in mathematical reasoning within visual contexts across various datasets. However, most existing multimodal math benchmarks are limited to single-visual contexts, which diverges from the multi-visual scenarios commonly encountered in real-world mathematical applications. To address this gap, we introduce MV-MATH: a meticulously curated dataset of 2,009 high-quality mathematical problems. Each problem integrates multiple images interleaved with text, derived from authentic K-12 scenarios, and enriched with detailed annotations. MV-MATH includes multiple-choice, free-form, and multi-step questions, covering 11 subject areas across 3 difficulty levels, and serves as a comprehensive and rigorous benchmark for assessing MLLMs' mathematical reasoning in multi-visual contexts. Through extensive experimentation, we observe that MLLMs encounter substantial challenges in multi-visual math tasks, with a considerable performance gap relative to human capabilities on MV-MATH. Furthermore, we analyze the performance and error patterns of various models, providing insights into MLLMs' mathematical reasoning capabilities within multi-visual settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。