让大模型同时看图解题,提升数学推理能力
MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
- 构建多模态数学数据集,融合图像与文字解题步骤
- 在30万道题目上训练,实现开源模型最佳多模态表现
- 适合需要图文结合解题的教育类应用
大型语言模型(LLM)在数学推理领域发展迅速,但多数开源模型仅关注纯文本数学任务,忽视了视觉输入的重要性。事实上,几何图、图表和函数图像等视觉信息对许多数学问题至关重要。为此,我们提出 extbf{MultiMath-7B},一个将数学与视觉能力融合的多模态大模型。该模型通过四阶段训练流程,重点强化视觉-语言对齐、视觉与数学指令微调,以及过程监督的强化学习。我们还构建了一个新颖、多样且全面的多模态数学数据集 extbf{MultiMath-300K},覆盖K-12阶段,包含图像描述和分步解题过程。MultiMath-7B 在现有开源多模态数学基准上达到领先性能,并在纯文本数学基准上同样表现优异。模型与数据集已开源:{ extcolor{blue}{ exttt{https://github.com/pengshuai-rin/MultiMath}}}。
原文摘要 · Abstract (English)
The rapid development of large language models (LLMs) has spurred extensive research into their domain-specific capabilities, particularly mathematical reasoning. However, most open-source LLMs focus solely on mathematical reasoning, neglecting the integration with visual injection, despite the fact that many mathematical tasks rely on visual inputs such as geometric diagrams, charts, and function plots. To fill this gap, we introduce \textbf{MultiMath-7B}, a multimodal large language model that bridges the gap between math and vision. \textbf{MultiMath-7B} is trained through a four-stage process, focusing on vision-language alignment, visual and math instruction-tuning, and process-supervised reinforcement learning. We also construct a novel, diverse and comprehensive multimodal mathematical dataset, \textbf{MultiMath-300K}, which spans K-12 levels with image captions and step-wise solutions. MultiMath-7B achieves state-of-the-art (SOTA) performance among open-source models on existing multimodal mathematical benchmarks and also excels on text-only mathematical benchmarks. Our model and dataset are available at {\textcolor{blue}{\url{https://github.com/pengshuai-rin/MultiMath}}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。