arXiv:2505.10557cs.CVcs.AI2025-05ACL被引 55

用代码做监督信号,提升模型理解数学图像的能力。

MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning

  • 以代码为桥梁实现图像与文本的精准对齐
  • 在六项指标上达到开源模型新纪录,几何题超越GPT-4o 8.9%
  • 适合需要高精度数学视觉推理的研究者与开发者

现有自然语言图像标注数据集多聚焦于自然场景,忽略数学图形中解题关键细节,制约了大模型在多模态数学推理中的发展。为此,本文提出利用代码作为跨模态对齐的监督信号,因代码能完整表达生成对应图形所需信息。我们采用模型在环的方法协同构建图像到代码模型FigCodifier及图像-代码数据集ImgCode-8.6M,为迄今最大规模。进一步,使用FigCodifier合成新数学图形,构建高质量多模态数学指令微调数据集MM-MathInstruct-3M。最终提出MathCoder-VL,先在ImgCode-8.6M上进行跨模态对齐训练,再在MM-MathInstruct-3M上微调,实现多模态数学问题求解。该模型在所有六项指标上达到开源模型新最优,尤其在MathVista几何子集上超越GPT-4o 8.9%、Claude 3.5 Sonnet 9.2%。相关数据集与模型将公开于https://github.com/mathllm/MathCoder。

原文摘要 · Abstract (English)

Natural language image-caption datasets, widely used for training Large Multimodal Models, mainly focus on natural scenarios and overlook the intricate details of mathematical figures that are critical for problem-solving, hindering the advancement of current LMMs in multimodal mathematical reasoning. To this end, we propose leveraging code as supervision for cross-modal alignment, since code inherently encodes all information needed to generate corresponding figures, establishing a precise connection between the two modalities. Specifically, we co-develop our image-to-code model and dataset with model-in-the-loop approach, resulting in an image-to-code model, FigCodifier and ImgCode-8.6M dataset, the largest image-code dataset to date. Furthermore, we utilize FigCodifier to synthesize novel mathematical figures and then construct MM-MathInstruct-3M, a high-quality multimodal math instruction fine-tuning dataset. Finally, we present MathCoder-VL, trained with ImgCode-8.6M for cross-modal alignment and subsequently fine-tuned on MM-MathInstruct-3M for multimodal math problem solving. Our model achieves a new open-source SOTA across all six metrics. Notably, it surpasses GPT-4o and Claude 3.5 Sonnet in the geometry problem-solving subset of MathVista, achieving improvements of 8.9% and 9.2%. The dataset and models will be released at https://github.com/mathllm/MathCoder.

多模态数学推理代码监督图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。