构建多样化数学视觉数据集,提升多模态模型解题能力
MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model
- 基于自建数据集MathVL进行监督微调,聚焦数学多模态推理
- 在2000题测试集上显著超越现有开源数学多模态模型
- 适合研究数学视觉推理、多模态大模型的开发者与学者
大语言模型在文本数学问题求解上表现突出,但现有数学专用多模态大模型主要关注几何问题,忽视其他数学领域丰富的视觉信息。且其几何数据多来自多样性与复杂性有限的公开数据集。为此,我们构建了名为MathVL的细调数据集,并基于不同参数规模的骨干网络,通过监督微调(SFT)开发了一系列专用于数学的多模态大模型——MathGLM-Vision。为全面评估其效果,我们在多个公开基准和自建的2000题MathVL-test上进行实验。结果表明,MathGLM-Vision在性能上显著优于部分现有模型,包括基础模型及开源数学多模态模型。这些发现凸显了多样化数据集对提升多模态大模型数学推理能力的重要性。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated significant capabilities in mathematical reasoning, particularly with text-based mathematical problems. However, current multi-modal large language models (MLLMs), especially those specialized in mathematics, tend to focus predominantly on solving geometric problems but ignore the diversity of visual information available in other areas of mathematics. Moreover, the geometric information for these specialized mathematical MLLMs is derived from several public datasets, which are typically limited in diversity and complexity. To address these limitations, we aim to construct a fine-tuning dataset named MathVL, and develop a series of specialized mathematical MLLMs termed MathGLM-Vision by conducting Supervised Fine-Tuning (SFT) on MathVL with various parameter-scale backbones. To extensively evaluate the effectiveness of MathGLM-Vision, we conduct experiments on several public benchmarks and our curated MathVL-test consisting of 2,000 problems. Experimental results demonstrate that MathGLM-Vision achieves significant improvements compared with some existing models, including backbone models and open-source mathematical MLLMs. These findings indicate the importance of diversity dataset in enhancing the mathematical reasoning abilities of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。