arXiv:2503.10291cs.CVcs.CL2025-03被引 132

用80亿参数模型提升多模态推理能力,效果显著且可复用。

VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

  • 基于自动数据管道构建40万条多模态推理过程数据,训练高效PRM模型。
  • 在7个基准上使最强模型InternVL2.5-78B提升5.9分,优于传统方法。
  • 提供带人工标注步骤的评估基准,助力多模态推理模型优化。

我们提出VisualPRM,一个拥有80亿参数的先进多模态过程奖励模型(PRM),通过最佳N次采样(Best-of-N)策略,显著提升多种规模和架构的多模态大语言模型(MLLMs)的推理能力。该模型在三类不同类型的MLLMs及四种不同规模上均表现优异,即使应用于强大的InternVL2.5-78B模型,在七个多模态推理基准上也实现了5.9分的性能提升。实验表明,VisualPRM在BoN评估中优于结果奖励模型和自一致方法。为支持多模态PRM训练,我们构建了包含40万条样本的VisualPRM400K多模态过程监督数据集,采用自动化数据流水线。针对评估需求,我们提出VisualProcessBench,一个带有人类标注步骤正确性的基准,用于衡量PRM检测多模态推理错误步骤的能力。相关模型、数据与基准已开源。

原文摘要 · Abstract (English)

We introduce VisualPRM, an advanced multimodal Process Reward Model (PRM) with 8B parameters, which improves the reasoning abilities of existing Multimodal Large Language Models (MLLMs) across different model scales and families with Best-of-N (BoN) evaluation strategies. Specifically, our model improves the reasoning performance of three types of MLLMs and four different model scales. Even when applied to the highly capable InternVL2.5-78B, it achieves a 5.9-point improvement across seven multimodal reasoning benchmarks. Experimental results show that our model exhibits superior performance compared to Outcome Reward Models and Self-Consistency during BoN evaluation. To facilitate the training of multimodal PRMs, we construct a multimodal process supervision dataset VisualPRM400K using an automated data pipeline. For the evaluation of multimodal PRMs, we propose VisualProcessBench, a benchmark with human-annotated step-wise correctness labels, to measure the abilities of PRMs to detect erroneous steps in multimodal reasoning tasks. We hope that our work can inspire more future research and contribute to the development of MLLMs. Our model, data, and benchmark are released in https://internvl.github.io/blog/2025-03-13-VisualPRM/.

多模态推理过程奖励扩散模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。