评测视频世界模型在复杂场景下的可信度,发现现有模型存在安全推理缺陷。
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

- 构建四类真实场景测试集,涵盖正常、约束、反事实和对抗情境。
- 7个模型均出现物理交互与危险指令抑制失败,视觉流畅但逻辑不稳。
- 适合关注机器人视觉建模可信性与安全性的研究者参考。
视频世界模型在机器人操作中日益重要,但现有基准多仅评估其在有效、可行且安全指令下的表现。本文提出RoboTrustBench,一个针对视频世界模型可信度的评测基准,涵盖四种场景:正常、约束敏感、反事实和对抗。该基准基于真实世界DROID数据集,包含1,207对专家验证的指令-图像样本,并采用六维评估协议,包含13项细粒度标准。通过人类与多模态大模型(MLLM)评估七种代表性视频世界模型,结果表明:当前模型虽能生成视觉连贯视频,但在约束推理、反事实定位、物理交互及危险指令抑制方面表现不佳。这说明仅依赖视觉质量与表面指令遵循不足以保证可信的机器人视频建模。
原文摘要 · Abstract (English)
Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instructions. We introduce RoboTrustBench, a benchmark for evaluating the trustworthiness of video world models under four scenarios: Normal, Constraint-Sensitive, Counterfactual, and Adversarial. Built from real-world DROID episodes, RoboTrustBench contains 1,207 expert-validated instruction-image pairs and a six-dimensional evaluation protocol with 13 fine-grained criteria. Evaluating seven representative video world models with human and MLLM assessment, we find that current models often generate visually coherent videos, but struggle with constraint reasoning, counterfactual grounding, physical interaction, and unsafe-instruction suppression. These results show that visual quality and surface-level instruction following are insufficient for trustworthy robotic video world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。