构建真实教育场景数学题数据集,评估大模型在复杂图像下的推理能力
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
- 基于2000张真实拍摄的教育图像构建新基准
- 发现现有模型在真实场景下表现大幅下降
- 适合关注多模态模型落地应用的研究者
多模态大语言模型在现有数学推理基准上表现优异,但这些基准多采用经过处理的干净图像,缺乏真实教育场景中的原始输入。为此,我们提出MathReal,一个精心构建的数据集,包含2000个由手持设备在真实场景中拍摄的数学问题图像。每道题均含文本与视觉元素。我们系统地将真实图像分为三类:图像质量退化、视角变化、无关内容干扰,并细分为14个子类。该数据集覆盖五个核心知识与能力类别,包含三种题型,分三个难度等级。为全面评估当前顶尖多模态模型在真实场景中的数学推理能力,我们设计六种实验设置,实现系统性分析。大量实验表明,现有模型在真实教育环境中面临严峻挑战。我们进一步分析其性能与错误模式,揭示其在识别、理解与推理方面的瓶颈,并指明未来改进方向。数据与代码见:https://github.com/junfeng0288/MathReal。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in visual mathematical reasoning across various existing benchmarks. However, these benchmarks are predominantly based on clean or processed multimodal inputs, without incorporating the images provided by real-world Kindergarten through 12th grade (K-12) educational users. To address this gap, we introduce MathReal, a meticulously curated dataset comprising 2,000 mathematical questions with images captured by handheld mobile devices in authentic scenarios. Each question is an image, containing the question text and visual element. We systematically classify the real images into three primary categories: image quality degradation, perspective variation, and irrelevant content interference, which are further delineated into 14 subcategories. Additionally, MathReal spans five core knowledge and ability categories, which encompass three question types and are divided into three difficulty levels. To comprehensively evaluate the multimodal mathematical reasoning abilities of state-of-the-art MLLMs in real-world scenarios, we design six experimental settings that enable a systematic analysis of their performance. Through extensive experimentation, we find that the problem-solving abilities of existing MLLMs are significantly challenged in realistic educational contexts. Based on this, we conduct a thorough analysis of their performance and error patterns, providing insights into their recognition, comprehension, and reasoning capabilities, and outlining directions for future improvements. Data and code: https://github.com/junfeng0288/MathReal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。