首个中文多模态数学数据集,用于评测和提升大模型的数学推理能力。
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
- 构建包含2.8万样本的中文多模态数学数据集,覆盖12个学段
- 发现当前顶尖多模态模型在该数据集上表现仍不理想
- 提出Math-LMM模型,三阶段训练显著提升数学推理能力
大型语言模型在数学推理任务中已取得良好进展,而以往研究主要基于纯文本数学数据集(如MATH、GSM8K)。近期有少数英文多模态数学数据集(如MATHVISTA和MATH-V)发布,用于评估大模型多模态能力。本文发布首个中文多模态数学数据集CMM-Math,包含基准与训练两部分,涵盖超过28,000个高质量样本,覆盖中国从小学到高中的12个年级,题型多样(如选择题、填空题等),并配有详细解题过程。题目中可能包含图像信息,使任务更具挑战性。通过全面分析,我们发现现有先进多模态模型在该数据集上仍面临困难,凸显其改进必要性。为此,我们提出Math-LMM模型,支持多图与文本混合输入。采用三阶段训练策略:基础预训练、基础微调与数学专项微调。大量实验表明,该模型在三个多模态数学数据集上均优于当前最先进方法。
原文摘要 · Abstract (English)
Large language models (LLMs) have obtained promising results in mathematical reasoning, which is a foundational skill for human intelligence. Most previous studies focus on improving and measuring the performance of LLMs based on textual math reasoning datasets (e.g., MATH, GSM8K). Recently, a few researchers have released English multimodal math datasets (e.g., MATHVISTA and MATH-V) to evaluate the effectiveness of large multimodal models (LMMs). In this paper, we release a Chinese multimodal math (CMM-Math) dataset, including benchmark and training parts, to evaluate and enhance the mathematical reasoning of LMMs. CMM-Math contains over 28,000 high-quality samples, featuring a variety of problem types (e.g., multiple-choice, fill-in-the-blank, and so on) with detailed solutions across 12 grade levels from elementary to high school in China. Specifically, the visual context may be present in the questions or opinions, which makes this dataset more challenging. Through comprehensive analysis, we discover that state-of-the-art LMMs on the CMM-Math dataset face challenges, emphasizing the necessity for further improvements in LMM development. We also propose a Multimodal Mathematical LMM (Math-LMM) to handle the problems with mixed input of multiple images and text segments. We train our model using three stages, including foundational pre-training, foundational fine-tuning, and mathematical fine-tuning. The extensive experiments indicate that our model effectively improves math reasoning performance by comparing it with the SOTA LMMs over three multimodal mathematical datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。