提出新评测框架,检验世界模型在真实机器人任务中的表现差距。
Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test
- 构建包含22项指标的综合评测体系,以人类偏好为基准验证生成质量。
- 模型在长期规划中仅达17.27分,物理一致性最高68.02分,显示时空推理不足。
- 首次用逆动态模型测试执行准确率,多数模型失败,仅WoW达40.74%成功率。
随着世界模型在具身AI中的兴起,越来越多研究尝试使用视频基础模型作为预测性世界模型,用于3D预测或交互生成等下游任务。但在探索这些任务前,视频基础模型仍面临两个关键问题:其生成泛化是否足以保持人类观察者眼中的感知保真度,以及是否足够鲁棒以作为真实世界具身智能体的通用先验。为此,我们提出具身图灵测试基准Wow-wo-val(Wow, wo, val),基于609个机器人操作数据,评估感知、规划、预测、泛化和执行五大核心能力。通过22项指标的综合评估协议,整体得分与人类偏好相关性超过0.93,建立可靠的真人图灵测试基础。在该基准上,模型在长时序规划中仅获17.27分,物理一致性最高68.02分,表明时空一致性和物理推理能力有限。针对逆动力学模型图灵测试,我们首次利用逆动态模型评估视频基础模型在真实世界中的执行准确性,结果显示多数模型成功率趋近0%,而WoW达到40.74%。这些发现揭示生成视频与真实世界间存在显著差距,凸显在具身AI中评测世界模型的紧迫性与必要性。
原文摘要 · Abstract (English)
As world models gain momentum in Embodied AI, an increasing number of works explore using video foundation models as predictive world models for downstream embodied tasks like 3D prediction or interactive generation. However, before exploring these downstream tasks, video foundation models still have two critical questions unanswered: (1) whether their generative generalization is sufficient to maintain perceptual fidelity in the eyes of human observers, and (2) whether they are robust enough to serve as a universal prior for real-world embodied agents. To provide a standardized framework for answering these questions, we introduce the Embodied Turing Test benchmark: WoW-World-Eval (Wow,wo,val). Building upon 609 robot manipulation data, Wow-wo-val examines five core abilities, including perception, planning, prediction, generalization, and execution. We propose a comprehensive evaluation protocol with 22 metrics to assess the models' generation ability, which achieves a high Pearson Correlation between the overall score and human preference (>0.93) and establishes a reliable foundation for the Human Turing Test. On Wow-wo-val, models achieve only 17.27 on long-horizon planning and at best 68.02 on physical consistency, indicating limited spatiotemporal consistency and physical reasoning. For the Inverse Dynamic Model Turing Test, we first use an IDM to evaluate the video foundation models' execution accuracy in the real world. However, most models collapse to $\approx$ 0% success, while WoW maintains a 40.74% success rate. These findings point to a noticeable gap between the generated videos and the real world, highlighting the urgency and necessity of benchmarking World Model in Embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。