arXiv:2606.28385cs.ROcs.AI2026-06

用多智能体框架精准诊断机器人视频生成的物理与逻辑错误。

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

论文配图:RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis
图 1 · 摘自论文原文
  • 分三阶段分析任务、定位维度、专家验证,实现可解释评估。
  • 相比零样本基线,描述准确率最高提升43点,时序对齐提升37点。
  • 适合研究世界模型或评估机器人视频生成的开发者使用。

近期机器人世界模型可生成用于具身预测与规划的合成视频,但评估困难:视觉逼真却常违反物理规律、时间不一致或任务逻辑错误。传统指标与单一视觉语言模型(VLM)无法泛化且缺乏诊断能力。我们提出RoboGaze,一种无需训练的多智能体VLM框架,实现结构化、可解释的生成视频评估。给定任务指令与视频,RoboGaze通过三阶段流程:任务-场景定位、维度特化专家路由、批判性验证,输出按新型6维度30类机器人专用分类的时序局部故障报告。为基准测试,我们构建了包含382个片段的人工验证数据集,涵盖模拟与真实世界多视角操作。评估8个开源及专有VLM主干,RoboGaze显著优于零样本基线,描述F1最高提升43点,时序对齐(F1 × IoU)最高提升37点,接近人类上限的85%。其批判验证器还缓解了标准VLM的‘狼来了’误报问题,使无错视频准确率从不足25%提升至超80%。RoboGaze提供了一种可扩展、高度可解释的诊断工具,适用于机器人世界模型的严格评估。

原文摘要 · Abstract (English)

Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is challenging: visually realistic outputs often violate physical laws, temporal consistency, or task logic, while conventional metrics and monolithic Vision-Language Model (VLM) judges fail to generalize or provide precise diagnostic value. We present RoboGaze, a training-free, multi-agent VLM framework that provides structured, interpretable evaluation for generated robot-manipulation videos. Given a task instruction and video, RoboGaze operates via a three-stage pipeline: task-scene grounding, dimension-specific specialist routing, and critic-based verification. It outputs temporally localized glitch reports categorized under a novel 6-dimension, 30-type robotics-specific taxonomy. To benchmark RoboGaze, we introduce a human-validated dataset of 382 clips spanning simulated and real-world multi-view manipulation. Evaluating eight open-source and proprietary VLM backbones, RoboGaze dramatically outperforms zero-shot baselines, improving description-F1 by up to +43 points and temporal alignment (F1 x IoU) by up to +37 points, closing approximately 85% of the gap to the human ceiling. Furthermore, its critic verifier mitigates the "cry-wolf" false-positive flaw of standard VLMs, lifting clean-clip accuracy from under 25% to over 80%. RoboGaze offers a scalable, highly interpretable diagnostic tool for the rigorous evaluation of robot world models.

机器人视频评估视觉语言模型世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。