用4D动态建模提升机器人操作的精准与协调能力
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

- 通过融合RGB、深度和光流构建统一4D生成模型
- 在真实双臂操作任务中达到顶尖性能,尤其擅长精细空间与时间控制
- 适合需要高精度动作预测的机器人视觉与决策研究者
开放世界中的机器人操作不仅需要理解场景外观,还需预判其3D结构在交互下的动态变化。我们提出,同步的RGB、深度与光流(即RGB-DF)构成一种物理基础表征,能捕捉场景的4D动态。相比2D像素视频,这种多模态协同使视觉外观与几何结构、时间运动对齐,更贴近机器人末端执行器的底层动作需求,缩小了世界预测与策略学习之间的差距。基于此,我们提出RynnWorld-4D,一个在统一扩散过程中从单个RGB-D图像和语言指令生成未来RGB帧、深度图与光流的生成模型。该模型采用三分支架构,结合跨模态注意力与逐帧3D RoPE,确保外观、几何与运动的一致演化。为大规模训练,我们构建了Rynn4DDataset 1.0,包含超过2.544亿帧的主视角人与机器人操作视频,附带高质量伪标签深度与光流。我们进一步提出RynnWorld-4D-Policy,一个逆动力学头,仅需一次前向传播即可利用模型内部4D表示输出机器人动作,实现闭环控制,无需昂贵的多步去噪。实验表明,RynnWorld-4D生成的4D预测在时空上高度一致,且RynnWorld-4D-Policy在真实世界灵巧双臂操作任务中表现领先,尤其在需空间精度与时间协调的任务中优势显著。
原文摘要 · Abstract (English)
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to the low-level end-effector actions demanded by robotic systems, thereby narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces temporally and spatially coherent 4D predictions, and that RynnWorld-4D-Policy achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。