arXiv:2503.12505cs.AIcs.CV2025-03ACL被引 16

构建多模态推理基准,评估模型找错与优化推理路径的能力

MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification

  • 设计三类评估范式,覆盖步骤正确性、答案聚合与推理搜索
  • 首次系统评测多模态过程奖励模型在复杂任务中的表现
  • 适合研究推理增强、强化学习与多模态模型的学者参考

推理是大语言模型应对复杂任务的核心能力,其中识别过程错误对提升该能力至关重要。近期提出的流程级奖励模型(PRMs)通过提供分步奖励,促进训练中的强化学习与数据生成,并在推理时引导模型走向正确步骤,从而提高推理准确性。然而,现有PRM基准多为文本导向,侧重错误检测,忽略了推理搜索等其他场景。为此,我们提出MPBench,一个综合性、多任务、多模态的基准,旨在系统评估PRMs在多种场景下的有效性。MPBench采用三种评估范式:(1) 步骤正确性,评估每一步推理的正确性;(2) 答案聚合,整合多个解法并选出最优结果;(3) 推理过程搜索,指导推理过程中寻找最优步骤序列。通过这些范式,MPBench实现全面评估,并为多模态PRMs的发展提供洞见。

原文摘要 · Abstract (English)

Reasoning is an essential capacity for large language models (LLMs) to address complex tasks, where the identification of process errors is vital for improving this ability. Recently, process-level reward models (PRMs) were proposed to provide step-wise rewards that facilitate reinforcement learning and data production during training and guide LLMs toward correct steps during inference, thereby improving reasoning accuracy. However, existing benchmarks of PRMs are text-based and focus on error detection, neglecting other scenarios like reasoning search. To address this gap, we introduce MPBench, a comprehensive, multi-task, multimodal benchmark designed to systematically assess the effectiveness of PRMs in diverse scenarios. MPBench employs three evaluation paradigms, each targeting a specific role of PRMs in the reasoning process: (1) Step Correctness, which assesses the correctness of each intermediate reasoning step; (2) Answer Aggregation, which aggregates multiple solutions and selects the best one; and (3) Reasoning Process Search, which guides the search for optimal reasoning steps during inference. Through these paradigms, MPBench makes comprehensive evaluations and provides insights into the development of multimodal PRMs.

多模态推理过程奖励基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。