arXiv:2602.08346cs.CV2026-02AAAI被引 8

首个面向图文推理过程的奖励模型评估基准,揭示现有模型在判断推理步骤上的不足。

What, Whether and How? Unveiling Process Reward Models for Thinking with Images Reasoning

  • 定义7类细粒度错误,构建图文推理轨迹评估框架。
  • 包含1206条人工标注轨迹,覆盖4大类16子类任务。
  • 发现当前视觉语言模型难以胜任过程奖励建模,存在明显偏差与位置敏感性。

大型视觉语言模型(LVLMs)在各类视觉任务中展现出卓越能力。在此基础上,

原文摘要 · Abstract (English)

The rapid advancement of Large Vision Language Models (LVLMs) has demonstrated excellent abilities in various visual tasks. Building upon these developments, the thinking with images paradigm has emerged, enabling models to dynamically edit and re-encode visual information at each reasoning step, mirroring human visual processing. However, this paradigm introduces significant challenges as diverse errors may occur during reasoning processes. This necessitates Process Reward Models (PRMs) for distinguishing positive and negative reasoning steps, yet existing benchmarks for PRMs are predominantly text-centric and lack comprehensive assessment under this paradigm. To address these gaps, this work introduces the first comprehensive benchmark specifically designed for evaluating PRMs under the thinking with images paradigm. Our main contributions are: (1) Through extensive analysis of reasoning trajectories and guided search experiments with PRMs, we define 7 fine-grained error types and demonstrate both the necessity for specialized PRMs and the potential for improvement. (2) We construct a comprehensive benchmark comprising 1,206 manually annotated thinking with images reasoning trajectories spanning 4 categories and 16 subcategories for fine-grained evaluation of PRMs. (3) Our experimental analysis reveals that current LVLMs fall short as effective PRMs, exhibiting limited capabilities in visual reasoning process evaluation with significant performance disparities across error types, positive evaluation bias, and sensitivity to reasoning step positions. These findings demonstrate the effectiveness of our benchmark and establish crucial foundations for advancing PRMs in LVLMs.

图文推理过程奖励评估基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。