arXiv:2410.04509cs.CL2024-10ACL被引 41

首个评估多模态模型数学错题识别能力的基准测试

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

  • 构建多模态错题检测任务,包含步骤定位与错误分类两个子任务
  • 使用2500个真实学生解题数据,GPT-4o表现仍比专家低10%
  • 适用于教育AI、数学推理模型评估与教学辅助系统研究者

随着多模态大语言模型(MLLMs)的发展,其在数学推理任务中的潜力日益凸显。现有数学评测主要关注问题求解能力,但缺乏对复杂场景如错误检测的评估。为此,本文正式提出多模态错误检测任务,并引入ErrorRadar——首个专门评估该能力的基准。该基准涵盖2,500个高质量、源自真实教育机构的学生解题数据,具有丰富标注信息(如题目类型、错误类别),支持误差步骤识别与错误分类两项子任务。通过大量实验对比开源与闭源代表性MLLMs,结果表明当前最佳模型GPT-4o仍比人类专家低约10%。这揭示了模型在复杂数学推理中错误理解与发现能力的重大挑战。

原文摘要 · Abstract (English)

As the field of Multimodal Large Language Models (MLLMs) continues to evolve, their potential to revolutionize artificial intelligence is particularly promising, especially in addressing mathematical reasoning tasks. Current mathematical benchmarks predominantly focus on evaluating MLLMs' problem-solving ability, yet there is a crucial gap in addressing more complex scenarios such as error detection, for enhancing reasoning capability in complicated settings. To fill this gap, we formally formulate the new task: multimodal error detection, and introduce ErrorRadar, the first benchmark designed to assess MLLMs' capabilities in such a task. ErrorRadar evaluates two sub-tasks: error step identification and error categorization, providing a comprehensive framework for evaluating MLLMs' complex mathematical reasoning ability. It consists of 2,500 high-quality multimodal K-12 mathematical problems, collected from real-world student interactions in an educational organization, with rigorous annotation and rich metadata such as problem type and error category. Through extensive experiments, we evaluated both open-source and closed-source representative MLLMs, benchmarking their performance against educational expert evaluators. Results indicate significant challenges still remain, as GPT-4o with best performance is still around 10% behind human evaluation.

数学推理多模态错误检测评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。