arXiv:2607.06559cs.RO2026-07被引 1

用4D动态建模提升机器人操作的精准与协调能力

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

论文配图:RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
图 1 · 摘自论文原文
  • 通过融合RGB、深度和光流构建统一4D生成模型
  • 在真实双臂操作任务中达到顶尖性能,尤其擅长精细空间与时间控制
  • 适合需要高精度动作预测的机器人视觉与决策研究者

开放世界中的机器人操作不仅需要理解场景外观,还需预判其3D结构在交互下的动态变化。我们提出,同步的RGB、深度与光流(即RGB-DF)构成一种物理基础表征,能捕捉场景的4D动态。相比2D像素视频,这种多模态协同使视觉外观与几何结构、时间运动对齐,更贴近机器人末端执行器的底层动作需求,缩小了世界预测与策略学习之间的差距。基于此,我们提出RynnWorld-4D,一个在统一扩散过程中从单个RGB-D图像和语言指令生成未来RGB帧、深度图与光流的生成模型。该模型采用三分支架构,结合跨模态注意力与逐帧3D RoPE,确保外观、几何与运动的一致演化。为大规模训练,我们构建了Rynn4DDataset 1.0,包含超过2.544亿帧的主视角人与机器人操作视频,附带高质量伪标签深度与光流。我们进一步提出RynnWorld-4D-Policy,一个逆动力学头,仅需一次前向传播即可利用模型内部4D表示输出机器人动作,实现闭环控制,无需昂贵的多步去噪。实验表明,RynnWorld-4D生成的4D预测在时空上高度一致,且RynnWorld-4D-Policy在真实世界灵巧双臂操作任务中表现领先,尤其在需空间精度与时间协调的任务中优势显著。

原文摘要 · Abstract (English)

Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to the low-level end-effector actions demanded by robotic systems, thereby narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.

4D建模机器人操作扩散模型多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。