评测视觉语言模型对机器人操作过程的理解能力,发现现有模型仍有明显短板。
RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

- 构建双维度评估框架:静态监控与动态推理,覆盖12类诊断问题。
- 基于58000+真实操作轨迹,涵盖260个任务,验证多模型过程理解缺陷。
- 提供可微调数据集,助力模型提升对操作阶段、动作进展的感知能力。
视觉语言模型(VLMs)正被探索用于机器人操作中的视觉评判、奖励生成和失败检测,这些角色隐含要求模型不仅能判断任务最终成败,还需理解操作过程的物理与时间演进。然而,现有评估无法检验VLM是否具备细粒度的过程理解能力。为此,本文提出RoboProcessBench,一个面向视觉语言机器人操作中过程感知理解的基准。该基准将此能力分解为静态监控与动态推理两个互补维度,设计了12类诊断问题,涵盖阶段、接触、运动、协调、原始动作进展、时间顺序、结果及原始动作级状态转换等。基准基于物理真实的执行轨迹构建,包含约58,000个问答对,覆盖260个操作任务,进一步划分为可用于后训练的ProcessData-SFT和用于评估的ProcessData-Eval。在ProcessData-Eval上的广泛评估显示,多种VLM在12类诊断任务中均表现出普遍局限,表明当前模型仍缺乏稳健的过程感知理解能力。但使用ProcessData-SFT进行后训练后,Qwen2.5-VL-7B和InternVL-3-8B模型在局部状态、运动、进展和原始动作提示上均展现出一致提升。结果表明,RoboProcessBench既是评估基准,也是可学习的监督信号,有助于发展具备过程监控与评估能力的VLM。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models to judge not only final task success, but also how a manipulation execution is physically and temporally progressing. However, existing evaluations fail to test whether VLMs possess fine-grained process understanding. To address this gap, we present RoboProcessBench, a benchmark for process-aware understanding in vision-language robotic manipulation. RoboProcessBench decomposes such capability into two complementary dimensions, \emph{static monitoring} and \emph{dynamic reasoning}, instantiated as 12 diagnostic question families covering phase, contact, motion, coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions. Built from physically grounded execution traces, the curated benchmark corpus ProcessData contains \textasciitilde 58k question-answer pairs across 260 manipulation tasks, which is further split into ProcessData-SFT and ProcessData-Eval for post-training and evaluation purposes. Extensive evaluation of various VLMs on ProcessData-Eval reveals broad limitations across 12 diagnostic task families, suggesting current models still lack robust process-aware understanding of manipulation executions. But with ProcessData-SFT, the post-trained \textit{Qwen2.5-VL-7B} and \textit{InternVL-3-8B} exhibit consistent gains on local state, motion, progress, and primitive-aware cues. These results demonstrate that RoboProcessBench serves as both an evaluation benchmark and a learnable supervision source for developing VLMs capable of monitoring and evaluating robotic manipulation processes. Project webpage: \href{https://processbench-2026.github.io/RoboProcessBench-Web/}{https://processbench-2026.github.io}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。