arXiv:2606.23296cs.RO2026-06

将运动学与物理动力学解耦,实现更真实可控的交互式世界模型。

IOI: Decoupling Kinematics and Physics for Interactive World Models

论文配图:IOI: Decoupling Kinematics and Physics for Interactive World Models
图 1 · 摘自论文原文
  • 用解析运动学先验生成精确轨迹,指导视频生成。
  • 在RoboTwin上实现最佳运动保真度和零样本泛化能力。
  • 适合需要高精度仿真与政策评估的研究者使用。

构建通用具身智能体需具备视觉逼真反馈和精准动作条件动态的交互环境。交互式世界模型通过模拟复杂动态来满足需求,但纯数据驱动方法因缺乏显式结构约束,难以保证控制对齐与物理合理视觉反馈。为此,我们提出IOI,一种融合解析运动学先验与学习物理动态的混合模型。不同于易产生时空漂移的数据驱动方法,IOI引入显式运动学引导,从动作序列计算前向运动学,生成准确运动轨迹。这些轨迹被渲染为同步的前、侧、顶正交投影图,无需外部相机标定。多视角运动学聚合与注入模块融合几何线索并注入视频生成器,提供几何一致引导。以确定性轨迹作为视频生成条件,建立解析模拟器与世界模型间的协同效应。将确定性运动解耦至运动学先验,使生成器专注建模随机物理交互。在RoboTwin基准测试中,IOI在运动保真度、分布外(OOD)泛化和策略评估方面均表现优异,达到当前最优仿真性能,并实现鲁棒的零样本泛化到未见的OOD任务。此外,IOI作为可靠策略评估工具,其成功率与真实物理模拟器高度一致。在真实平台上的实验表明,基于IOI合成数据训练的策略与基于遥操作演示训练的策略相当,验证了其在具身策略学习中的实际价值。

原文摘要 · Abstract (English)

Developing generalist embodied agents requires interactive environments providing visually realistic feedback and accurate action-conditioned dynamics. Interactive world models address this by simulating such complex dynamics. However, purely data-driven methods struggle to ensure precise control alignment and physically plausible visual feedback due to a lack of explicit structural constraints. To address this, we propose IOI, a hybrid interactive world model integrating analytical kinematic priors with learned physical dynamics. Unlike data-driven approaches prone to spatiotemporal drift, IOI introduces explicit kinematic guidance, computing forward kinematics from action sequences for accurate motion trajectories. These trajectories are rendered into synchronized front, side, and top orthographic projections, eliminating the need for extrinsic camera calibration. A Multi-view Kinematic Aggregation and Injection module fuses these geometric cues and injects them into the video generator, providing geometry-consistent guidance. Conditioning video generation on these deterministic trajectories establishes a synergy between the analytical simulator and the world model. Decoupling deterministic motion into the kinematic prior frees the generator to model stochastic physical interactions. Experiments on the RoboTwin benchmark validate IOI across kinematic fidelity, out-of-distribution (OOD) generalization, and policy evaluation. IOI achieves state-of-the-art simulation performance and robust zero-shot generalization to unseen OOD tasks. Furthermore, IOI serves as a reliable policy evaluator, yielding success rates closely aligning with ground-truth physics simulators. On real-world platforms, policies trained on IOI-synthesized data match those trained on teleoperation demonstrations, solidifying its practical value for embodied policy learning.

世界模型运动学仿真具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。