用4D动态建模实现机器人与环境的精准交互仿真。
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
- 基于运动学控制生成精确4D机器人轨迹,解耦动作与环境响应。
- 通过点云信号驱动生成模型,同步输出RGB与点云序列,还原复杂动态反应。
- 首个具备零样本迁移潜力的4D机器人仿真框架,适合构建高保真模拟环境。
机器人-世界交互仿真是具身智能的核心。现有方法多局限于2D空间或依赖静态环境线索,忽视了交互本质上是4D时空事件。为此,我们提出Kinema4D,一种动作条件的4D生成式机器人仿真器,将交互解耦为:i) 基于URDF的3D机器人运动学驱动,生成精确4D控制轨迹;ii) 将4D轨迹投影为时空点云信号,驱动生成模型合成复杂环境的响应动态,产出同步的RGB/点云序列。为支持训练,我们构建了大规模数据集Robo4D-200k,包含201,426条高质量4D标注的交互片段。大量实验表明,该方法能有效模拟物理合理、几何一致且与具体机器人无关的交互,真实还原多样现实动态。首次展现出零样本迁移潜力,为下一代具身仿真提供高保真基础。
原文摘要 · Abstract (English)
Simulating robot-world interactions is a cornerstone of Embodied AI. Recently, a few works have shown promise in leveraging video generations to transcend the rigid visual/physical constraints of traditional simulators. However, they primarily operate in 2D space or are guided by static environmental cues, ignoring the fundamental reality that robot-world interactions are inherently 4D spatiotemporal events that require precise interactive modeling. To restore this 4D essence while ensuring the precise robot control, we introduce Kinema4D, a new action-conditioned 4D generative robotic simulator that disentangles the robot-world interaction into: i) Precise 4D representation of robot controls: we drive a URDF-based 3D robot via kinematics, producing a precise 4D robot control trajectory. ii) Generative 4D modeling of environmental reactions: we project the 4D robot trajectory into a pointmap as a spatiotemporal visual signal, controlling the generative model to synthesize complex environments' reactive dynamics into synchronized RGB/pointmap sequences. To facilitate training, we curated a large-scale dataset called Robo4D-200k, comprising 201,426 robot interaction episodes with high-quality 4D annotations. Extensive experiments demonstrate that our method effectively simulates physically-plausible, geometry-consistent, and embodiment-agnostic interactions that faithfully mirror diverse real-world dynamics. For the first time, it shows potential zero-shot transfer capability, providing a high-fidelity foundation for advancing next-generation embodied simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。