构建3D对齐的机器人世界模型评估基准,精准区分生成与重建误差。
RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

- 基于统一3D重建流程,分离生成与重建误差,提升评估可靠性。
- 包含50项任务、25,000段多视角视频,覆盖4类场景,50个指标分4层级。
- 引入RoboPhyscore,与人类评价高度一致,适合评估模型执行能力。
视频世界模型在具身AI中日益作为数据引擎、行为规划器和模拟器使用,但现有具身世界模型(EWM)基准缺乏统一的3D对齐评估协议,难以判断生成轨迹是否保持真实3D场景状态或可执行。本文提出RoboPhys-3D,一个基于RoboTwin 2.0构建的3D对齐EWM基准,涵盖4种任务范式下的50项操作任务,共5,000个回合和25,000段多视角真值视频。其核心特性在于:生成视频与真值视频均通过相同3D重建流程处理,从而可区分重建误差与生成误差。该基准整合50个互补指标,划分为18个子维度,覆盖像素级保真度、3D几何一致性、状态理解与任务完成度四个层次。我们进一步提出平均全分(Average Full Score),综合所有50项指标;以及RoboPhyscore,聚焦与任务成功率强相关的指标进行加权。在四种代表性视频世界模型中,Cosmos 3取得最高RoboPhyscore(0.6330,达真值的92.7%)。而状态与执行层面的指标揭示了感知与视觉语言模型忽视的重大缺陷。RoboPhyscore与人类评估高度一致(皮尔逊相关r=0.9761,斯皮尔曼等级相关ρ=0.8962),凸显基于执行的接地评估对衡量EWM能力的重要性。
原文摘要 · Abstract (English)
Video world models increasingly serve as data engines, action planners, and simulators for embodied AI, but conventional embodied world model (EWM) benchmarks lack a unified 3D-grounded protocol for establishing whether generated rollouts preserve the underlying 3D scene state or translate into executable actions. We introduce RoboPhys-3D, a 3D-grounded EWM benchmark built on RoboTwin 2.0, covering 50 manipulation tasks across four regimes, with 5,000 episodes and 25,000 multi-view ground-truth videos. A defining feature of RoboPhys-3D is that generated and ground-truth videos are processed through the same 3D reconstruction pipeline, enabling reconstruction-induced error to be distinguished from generation-induced error. The RoboPhys-3D benchmark organizes 50 complementary metrics into 18 sub-dimensions across four levels: pixel-level fidelity, 3D geometry consistency, state-level understanding, and task-level completeness. We further introduce Average Full Score, a hierarchical score averaging all 50 metrics for comprehensive evaluation, and RoboPhyscore, a compact task-aligned score averaging the metrics most strongly correlated with task success. Among the four representative video world models, Cosmos 3 achieves the highest RoboPhyscore (0.6330, 92.7% of ground truth), while state- and execution-grounded metrics reveal substantial failures that perceptual and vision-language model-based judgments fail to capture. RoboPhyscore further exhibits strong agreement with human evaluation (Pearson r = 0.9761 and Spearman \r{ho} = 0.8962), demonstrating the importance of grounded, execution-aware evaluation for EWM capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。