arXiv:2509.23250cs.AIcs.CV2025-09被引 9

用更准的监督信号提升多模态模型推理可靠性,尤其在测试时增强表现。

Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned

  • 融合MCTS与强视觉语言模型生成更准确的步骤标签
  • 感知级监督显著提升多模态推理中视觉定位错误的检测能力
  • 小模型也能媲美大模型,适合资源有限场景

过程奖励模型(PRMs)提供步骤级监督,提升大语言模型推理的可靠性。尽管文本领域已有广泛研究,其在视觉语言模型(VLMs)中的应用仍有限。现有视觉语言过程奖励模型(VL-PRMs)依赖蒙特卡洛树搜索(MCTS)构建数据,常产生噪声监督信号,限制任务泛化能力。本文通过探索数据构建、训练与测试时扩展(TTS)策略,系统揭示了VL-PRMs的设计空间。首先提出混合数据合成框架,结合MCTS与强VLM判断,生成更精准的步骤标签;其次引入感知聚焦监督,使PRM能显式检测推理中视觉定位阶段的错误;最后系统评估多种测试时扩展策略,证明所提PRMs可有效引导VLM获得更准确解。在五个多样化多模态基准(MMMU、PuzzleVQA、AlgoPuzzleVQA、MathVista、MathVision)上的实验表明:(i) 将VL-PRMs作为结果奖励模型(ORMs)在测试时扩展中优于步骤选择引导;(ii) 更小的VL-PRMs可匹配甚至超越更大模型的错误检测能力;(iii) VL-PRMs可挖掘更强VLM骨干的潜在推理能力;(iv) 感知级监督在测试时扩展中带来显著性能提升;(v) 即使未在高级数学推理数据集上训练,不同策略在这些数据集上的测试时扩展表现仍持续改善。

原文摘要 · Abstract (English)

Process Reward Models (PRMs) provide step-level supervision that improves the reliability of reasoning in large language models. While PRMs have been extensively studied in text-based domains, their extension to Vision Language Models (VLMs) remains limited. Existing Vision-Language PRMs (VL-PRMs) rely on Monte Carlo Tree Search (MCTS) for data construction, which can often produce noisy supervision signals and limit generalization across tasks. In this work, we aim to elucidate the design space of VL-PRMs by exploring diverse strategies for dataset construction, training, and test-time scaling. First, we introduce a hybrid data synthesis framework that combines MCTS with judgments from a strong VLM, producing more accurate step-level labels. Second, we propose perception-focused supervision, enabling our PRM to explicitly detect errors at the visual grounding stage of reasoning. Third, we systematically evaluate multiple test-time scaling strategies, showing that our PRMs can reliably guide VLMs toward more accurate solutions. Our experiments covering five diverse multimodal benchmarks (MMMU, PuzzleVQA, AlgoPuzzleVQA, MathVista, and MathVision) reveal several key insights: (i) VL-PRMs when used as Outcome Reward Models (ORMs) during test-time scaling (TTS) can outperform VL-PRM guided process step selection, (ii) smaller VL-PRMs can match or even surpass larger ones in detecting process errors, (iii) VL-PRMs uncover latent reasoning abilities in stronger VLM backbones, (iv) perception-level supervision leads to significant gains in test-time scaling, and (v) TTS performance of different policies improve on advanced math reasoning datasets despite not training VL-PRMs on such datasets. We hope our work will motivate further research and support the advancement of VLMs.

多模态推理奖励模型测试时扩展视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。