评测视频世界模型在机器人操作中的可执行性,发现视觉真实不等于物理可行。
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

- 将生成的操控视频转为机器人动作序列,在物理仿真中验证执行效果。
- 评估结果显示视觉逼真度与实际可操作性常不一致,尤其在复杂长程任务中。
- 适合关注机器人具身智能、世界模型落地的研究者和开发者。
近年来大规模视频世界模型在未来预测方面取得进展,有望以生成视频作为机器人学习的可扩展监督信号。然而对于具身操控任务,仅具备感知真实性不足:生成的交互必须在物理上合理且可被机器人执行。现有基准虽评估了视觉质量与物理合理性,但未系统检验预测行为能否转化为可执行动作以完成操纵任务。本文提出RoboWM-Bench,一个以操纵为核心、基于具身化评估的视频世界模型评测基准。该基准通过真实场景重建与多样化的操控任务,将生成的人手及机器人操控视频转换为具身动作序列,并在物理仿真环境中进行执行验证。实验表明,视觉逼真度与具身可执行性并不总是一致。分析揭示空间推理、接触预测及非物理几何失真等因素显著影响执行表现,尤其在复杂和长时程交互中。这些发现为当前模型能力提供了更精细的认知,凸显了具身感知评估对机器人操控中物理合理世界建模的重要性。
原文摘要 · Abstract (English)
Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for robot learning. However, for embodied manipulation, perceptual realism alone is not sufficient: generated interactions must also be physically consistent and executable by robotic agents. Existing benchmarks provide valuable assessments of visual quality and physical plausibility, but they do not systematically evaluate whether predicted behaviors can be translated into executable actions that complete manipulation tasks. We introduce RoboWM-Bench, a manipulation-centric benchmark for embodiment-grounded evaluation of video world models. RoboWM-Bench converts generated human-hand and robotic manipulation videos into embodied action sequences and validates them through execution in physically grounded simulation environments. Built on real-to-sim scene reconstruction and diverse manipulation tasks, RoboWM-Bench enables standardized, reproducible, and scalable evaluation of physical executability. Using RoboWM-Bench, we evaluate state-of-the-art video world models and observe that visual plausibility and embodied executability are not always aligned. Our analysis highlights several recurring factors that affect execution performance, including spatial reasoning, contact prediction, and non-physical geometric distortions, particularly in complex and long-horizon interactions. These findings provide a more fine-grained view of current model capabilities and underscore the value of embodiment-aware evaluation for guiding physically grounded world modeling in robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。