不依赖逆动力学模型,用视觉直接提取机械臂位姿来评估物理一致性。
KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding

- 用视觉基础模型直接从视频帧提取6维末端执行器位姿,跳过易出错的逆动力学模型。
- 在ManiSkill3上测试4类任务,发现生成能力随复杂度呈非线性增长。
- 适合研究具身世界模型、机器人动作生成与物理仿真评估的学者。
评估具身世界模型(EWMs)的物理一致性是一项关键挑战。尽管闭环评估通过模拟器回放能更真实地检验物理合理性,但现有框架几乎全部依赖逆动力学模型(IDMs)进行动作提取。由于从2D像素空间到3D运动学空间映射复杂,训练外的数据会使学习到的IDMs变得脆弱,导致在包含新物体和场景的生成视频中动作提取不可靠,造成世界模型误差与提取器错误之间的归属模糊。为减少这种模糊性,我们提出KineBench——一种无需逆动力学模型的闭环评估基准,基于显式的运动学定位流程。给定生成视频,KineBench利用级联视觉基础模型直接从单帧中提取6维末端执行器位姿,并在物理模拟器中执行以进行闭环验证。除了基于任务完成的评估外,还引入两个经典3D运动学指标:谱弧长(SPARC)和丸山可操作性指数,从机器人视角刻画轨迹平滑性和运动可行性。该基准在ManiSkill3的20个多样化操作任务上构建,涵盖四个递进评估套件:基础执行、任务迁移、视觉分布外泛化以及复杂度条件下的缩放。对前沿模型的评估揭示了具身视频生成存在任务复杂度受限的非线性缩放现象,为未来数据扩展策略提供了实证指导。
原文摘要 · Abstract (English)
Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-loop evaluation via simulator rollouts offers a more faithful assessment of physical plausibility than open-loop alternatives, existing frameworks almost exclusively rely on Inverse Dynamics Models(IDMs) for action extraction. Due to the intricate mapping from 2D pixel space to 3D kinematic space, the learned IDMs can be brittle to data outside their training distribution, resulting in unreliable action extraction from the generated videos with novel objects and scenarios. This creates an unavoidable attribution ambiguity between world model inaccuracies and extractor errors. To reduce this ambiguity, we present KineBench, an IDM-free closed-loop benchmark for EWMs, built upon an explicit kinematic grounding pipeline. Given a generated video, KineBench employs cascaded visual foundation models to directly extract 6D end-effector poses from individual frames, which are then executed in a physics simulator for closed-loop validation. Beyond execution-based task success, KineBench incorporates two classical 3D kinematic metrics--Spectral Arc Length (SPARC) and the Maruyama Manipulability Index--to characterize trajectory smoothness and kinematic feasibility from a robot-centric perspective. Built on 20 diverse manipulation tasks in ManiSkill3, KineBench evaluates EWMs across four progressive suites: basic execution, task transfer, visual out-of-distribution generalization, and complexity-conditioned scaling. Evaluation across frontier models reveals task-complexity-bounded nonlinear scaling in embodied video generation, providing empirical guidance for future data-scaling strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。