arXiv:2503.07478cs.CV2025-03ICCV被引 21

构建首个覆盖推理全过程的视觉语言奖励模型评测基准。

VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models

  • 设计12项任务,分过程理解、结果判断、批评生成三类评估
  • 涵盖12634个问题,挑战模型在多图理解等场景下的表现
  • 支持开源与闭源模型对比,揭示当前先进模型仍有明显短板

尽管大型视觉语言模型(LVLM)在多模态任务中表现优异,但推理过程中的偏差可能导致错误。近年来,奖励模型(RMs)在推理中日益关键:过程型RM评估每一步推理,结果型RM关注最终结论,批判型RM则对整个推理过程进行错误分析并提出修正。然而,现有视觉语言奖励模型(VLRM)的评测基准通常仅评估单一能力(如区分两个答案),限制了全面评估和模型发展。为此,我们提出一个全面且具有挑战性的基准VLRMBench,包含12,634个问题。该基准基于三种不同数据集构建,涵盖数学推理、幻觉理解及多图像理解。设计12项任务,覆盖过程理解、结果判断与批评生成三大方面。在21个开源模型和5个先进闭源模型上开展广泛实验,结果显示其挑战性显著:例如在二分类任务‘Forecasting Future’中,GPT-4o准确率仅为76.0%。我们还进行了深入分析,为未来VLRM的发展提供重要启示。代码与数据集将公开于https://github.com/JCruan519/VLRMBench。

原文摘要 · Abstract (English)

Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Recently, reward models (RMs) have become increasingly pivotal in the reasoning process. Specifically, process RMs evaluate each reasoning step, outcome RMs focus on the assessment of reasoning results, and critique RMs perform error analysis on the entire reasoning process, followed by corrections. However, existing benchmarks for vision-language RMs (VLRMs) typically assess only a single aspect of their capabilities (e.g., distinguishing between two answers), thus limiting the all-round evaluation and restricting the development of RMs in the visual-language domain. To address this gap, we propose a comprehensive and challenging benchmark, dubbed as VLRMBench, encompassing 12,634 questions. VLRMBench is constructed based on three distinct types of datasets, covering mathematical reasoning, hallucination understanding, and multi-image understanding. We design 12 tasks across three major categories, focusing on evaluating VLRMs in the aspects of process understanding, outcome judgment, and critique generation. Extensive experiments are conducted on 21 open-source models and 5 advanced closed-source models, highlighting the challenges posed by VLRMBench. For instance, in the `Forecasting Future', a binary classification task, the advanced GPT-4o achieves only a 76.0% accuracy. Additionally, we perform comprehensive analytical studies, offering valuable insights for the future development of VLRMs. We anticipate that VLRMBench will serve as a pivotal benchmark in advancing VLRMs. Code and datasets will be available at https://github.com/JCruan519/VLRMBench.

奖励模型多模态评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。