提出新评估框架,精准检测动作条件世界模型的物理仿真真实性
WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation

- 设计可观测模拟器契约,要求动作必须引发真实运动
- 在18000+实例上发现多模型存在动作实现退化与交互错误
- 适合关注仿真真实性的机器人规划与训练研究者使用
动作条件世界模型(ACWMs)有望为具身智能提供可扩展的预测模拟器,用于规划、策略评估和数据生成。其实现依赖于精确的动作-状态转移,而非仅视觉逼真。然而现有评估多关注视觉质量或任务结果,缺乏对模拟器保真度的直接检验。为此,本文提出可观测模拟器契约:给定动作应引发对应主体运动,环境响应须基于实际运动。据此构建WorldSimProbe,包含五类控制实验:局部控制敏感性、全局轨迹变化、多样动作源、交互根基性及动力学特性。针对各实验设计评估器,检验模拟器校准度、动作到运动的密集对应、虚假交互及基础动力学表现。在RoboTwin、ManiSkill和LIBERO上对六种开源ACWMs进行超过18,000次测试,发现控制变化下动作实现性能系统性下降,交互根基性和动力学存在结构性失败,且评估信号与人类判断和下游任务表现一致。该能力导向框架为超越粗粒度任务评估提供了透明、标准化的模拟器保真度诊断范式。
原文摘要 · Abstract (English)
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs. Yet their applicability remains difficult to establish because prevailing evaluations emphasize visual quality, task outcomes, or coarse rollout-level responsiveness without directly testing simulator fidelity. To address this gap, we evaluate ACWMs through the observable capabilities expected of physical simulators. Accordingly, we formalize Observable Simulator Contract, a minimal contract that any action-conditioned physical simulator should satisfy: supplied actions must induce corresponding agent motion, and environment responses must be grounded in that realized motion. To operationalize this contract, we introduce WorldSimProbe, comprising five controlled suites spanning local control sensitivity, global trajectory variation, source-diverse actions, interaction grounding, and dynamics. Suite-specific evaluators assess simulator-relative calibration, dense action-to-motion correspondence, false-interaction grounding, and primitive-level dynamics. We evaluate six open-source ACWMs on more than 18,000 instances across RoboTwin, ManiSkill, and LIBERO. World-SimProbe reveals systematic action-realization degradation across control variation, structured failures in interaction grounding and dynamics, and benchmark signals consistent with human judgments and downstream outcomes. Together, this capability-based framework provides a transparent, and standardized paradigm for diagnosing ACWM simulator fidelity beyond coarse, task-directed evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。