arXiv:2508.10433cs.AIcs.CV2025-08被引 31

构建数学知识体系与强化学习训练,提升多模态模型解题能力

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

  • 设计五层数学知识体系,覆盖491个知识点和1819条基础原理
  • 生成7种难度渐进的题目变体,构建挑战性数据集以提升模型泛化能力
  • 采用两阶段强化学习框架,实现从知识对齐到难度递进的渐进式训练

多模态大语言模型在多项任务中表现优异,但在复杂数学推理上仍显不足。现有研究多聚焦于数据集构建与方法优化,忽视了知识驱动设计与模型为中心的数据空间建模。本文提出We-Math 2.0,一个融合结构化数学知识系统、模型中心数据空间建模与基于强化学习的训练范式的统一系统,全面增强多模态模型的数学推理能力。主要贡献包括:(1) 构建五级层次化数学知识体系,涵盖491个知识点与1819条基础原理;(2) 设计MathBook-Standard数据集,通过双重扩展确保概念覆盖面与灵活性;定义三维难度空间,每道题生成7个渐进变体,构建MathBook-Pro挑战数据集;(3) 提出两阶段强化学习框架:冷启动微调使模型对齐知识导向的思维链;渐进对齐强化学习结合平均奖励学习与动态数据调度,在不同难度层级实现渐进对齐;(4) 提出MathBookEval基准,覆盖全部491个知识点,包含多样化的推理步骤分布。实验表明,MathBook-RL在四个主流基准上表现优于现有基线,在MathBookEval上取得优异成绩,显示出良好的数学推理泛化能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various tasks, but still struggle with complex mathematical reasoning. Existing research primarily focuses on dataset construction and method optimization, often overlooking two critical aspects: comprehensive knowledge-driven design and model-centric data space modeling. In this paper, we introduce We-Math 2.0, a unified system that integrates a structured mathematical knowledge system, model-centric data space modeling, and a reinforcement learning (RL)-based training paradigm to comprehensively enhance the mathematical reasoning abilities of MLLMs. The key contributions of We-Math 2.0 are fourfold: (1) MathBook Knowledge System: We construct a five-level hierarchical system encompassing 491 knowledge points and 1,819 fundamental principles. (2) MathBook-Standard & Pro: We develop MathBook-Standard, a dataset that ensures broad conceptual coverage and flexibility through dual expansion. Additionally, we define a three-dimensional difficulty space and generate 7 progressive variants per problem to build MathBook-Pro, a challenging dataset for robust training. (3) MathBook-RL: We propose a two-stage RL framework comprising: (i) Cold-Start Fine-tuning, which aligns the model with knowledge-oriented chain-of-thought reasoning; and (ii) Progressive Alignment RL, leveraging average-reward learning and dynamic data scheduling to achieve progressive alignment across difficulty levels. (4) MathBookEval: We introduce a comprehensive benchmark covering all 491 knowledge points with diverse reasoning step distributions. Experimental results show that MathBook-RL performs competitively with existing baselines on four widely-used benchmarks and achieves strong results on MathBookEval, suggesting promising generalization in mathematical reasoning.

数学推理强化学习多模态模型知识系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。