让大模型像人一样画图解题,突破数学推理的视觉瓶颈
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
- 构建内生视觉思维链,分两阶段训练模型生成并编辑高精度数学图示
- 在3000题基准上相比强基线提升86%相对性能,通用性优异
- 适合需要复杂可视化推理的研究者与教育科技开发者
尽管大型语言模型在文本推理中表现优异,但在依赖视觉辅助的数学领域(如几何)仍面临挑战。现有视觉思维链(VCoT)方法受限于固定外部工具或无法生成高质量、时机精准的图示。为此,我们提出MathCanvas框架,赋予统一的大规模多模态模型内在的视觉思维链能力。该框架包含两个阶段:首先,在1520万对数据集上预训练,包括1000万张图文对(MathCanvas-Imagen)和520万步进式编辑轨迹(MathCanvas-Edit),以掌握图示生成与编辑;其次,在21.9万例交错式视觉-文本推理路径数据集(MathCanvas-Instruct)上微调,学习何时及如何使用视觉辅助。为严格评估,我们构建了3000题的MathCanvas-Bench基准,要求模型输出交错式视觉-文本解法。基于该框架训练的BAGEL-Canvas模型,在MathCanvas-Bench上相较强基线实现86%的相对提升,并展现出对其他公开数学基准的良好泛化能力。本工作提供完整工具包、数据集与基准,推动大模型实现类人级视觉辅助推理。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have excelled in textual reasoning, they struggle with mathematical domains like geometry that intrinsically rely on visual aids. Existing approaches to Visual Chain-of-Thought (VCoT) are often limited by rigid external tools or fail to generate the high-fidelity, strategically-timed diagrams necessary for complex problem-solving. To bridge this gap, we introduce MathCanvas, a comprehensive framework designed to endow unified Large Multimodal Models (LMMs) with intrinsic VCoT capabilities for mathematics. Our approach consists of two phases. First, a Visual Manipulation stage pre-trains the model on a novel 15.2M-pair corpus, comprising 10M caption-to-diagram pairs (MathCanvas-Imagen) and 5.2M step-by-step editing trajectories (MathCanvas-Edit), to master diagram generation and editing. Second, a Strategic Visual-Aided Reasoning stage fine-tunes the model on MathCanvas-Instruct, a new 219K-example dataset of interleaved visual-textual reasoning paths, teaching it when and how to leverage visual aids. To facilitate rigorous evaluation, we introduce MathCanvas-Bench, a challenging benchmark with 3K problems that require models to produce interleaved visual-textual solutions. Our model, BAGEL-Canvas, trained under this framework, achieves an 86% relative improvement over strong LMM baselines on MathCanvas-Bench, demonstrating excellent generalization to other public math benchmarks. Our work provides a complete toolkit-framework, datasets, and benchmark-to unlock complex, human-like visual-aided reasoning in LMMs. Project Page: https://mathcanvas.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。