arXiv:2501.04686cs.CLcs.AI2025-01NeurIPS被引 36

用过程奖励模型提升多模态数学推理能力,效果超越GPT-4o。

Unlocking Multimodal Mathematical Reasoning via Process Reward Model

  • 构建多模态思维链数据集MMathCoT-1M,训练更强的数学推理模型。
  • 提出自动过程标注方法,实现逻辑与视觉一致性兼顾的监督信号。
  • 设计新型多模态强化学习算法,性能优于GPT-4o和Gemma3-12B。

过程奖励模型(PRMs)在通过测试时扩展(TTS)提升大语言模型(LLMs)数学推理能力方面展现出潜力,但其在多模态推理中的应用尚未被探索。本文首次尝试解锁PRMs在多模态数学推理中的潜力,识别三大挑战:高质量推理数据稀缺限制了基础多模态大模型(MLLMs)的能力,进而制约TTS与强化学习(RL)上限;多模态场景中缺乏自动化过程标注方法;单模态强化学习中的奖励作弊问题可能延伸至多模态场景。为此,我们提出三阶段展开式多模态过程监督辅助训练框架URSA。首先构建高质量大规模多模态思维链(CoT)数据集MMathCoT-1M,训练更强的数学推理基础模型URSA-8B。随后通过自动流程合成过程监督数据,强调逻辑正确性与感知一致性,引入DualMath-1.1M以支持URSA-8B-RM训练。最后提出过程监督组相对策略优化(PS-GRPO),开创性地实现多模态PRM辅助在线强化学习。应用PS-GRPO后,URSA-8B-PS-GRPO在6个基准上平均优于Gemma3-12B 8.4%、GPT-4o 2.7%。代码、数据与检查点见https://github.com/URSA-MATH。

原文摘要 · Abstract (English)

Process Reward Models (PRMs) have shown promise in enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) through Test-Time Scaling (TTS). However, their integration into multimodal reasoning remains largely unexplored. In this work, we take the first step toward unlocking the potential of PRMs in multimodal mathematical reasoning. We identify three key challenges: (1) the scarcity of high-quality reasoning data constrains the capabilities of foundation Multimodal Large Language Models (MLLMs), which imposes further limitations on the upper bounds of TTS and reinforcement learning (RL); (2) a lack of automated methods for process labeling within multimodal contexts persists; (3) the employment of process rewards in unimodal RL faces issues like reward hacking, which may extend to multimodal scenarios. To address these issues, we introduce URSA, a three-stage Unfolding multimodal Process-Supervision Aided training framework. We first construct MMathCoT-1M, a high-quality large-scale multimodal Chain-of-Thought (CoT) reasoning dataset, to build a stronger math reasoning foundation MLLM, URSA-8B. Subsequently, we go through an automatic process to synthesize process supervision data, which emphasizes both logical correctness and perceptual consistency. We introduce DualMath-1.1M to facilitate the training of URSA-8B-RM. Finally, we propose Process-Supervised Group-Relative-Policy-Optimization (PS-GRPO), pioneering a multimodal PRM-aided online RL method that outperforms vanilla GRPO. With PS-GRPO application, URSA-8B-PS-GRPO outperforms Gemma3-12B and GPT-4o by 8.4% and 2.7% on average across 6 benchmarks. Code, data and checkpoint can be found at https://github.com/URSA-MATH.

多模态数学推理强化学习过程奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。