arXiv:2504.20995cs.CVcs.RO2025-04被引 84

学习4D动态世界模型,让机器人能预判环境变化。

TesserAct: Learning 4D Embodied World Models

  • 用RGB-DN视频训练,融合颜色、深度和法向信息建模
  • 生成的4D场景在时空上保持一致,支持新视角合成
  • 适合需要精准环境预测的机器人操控与策略学习

本文提出一种高效学习4D具身世界模型的方法,能够预测具身智能体行动下3D场景随时间的动态演变,保证空间与时间的一致性。通过在RGB-DN(RGB、深度、法向)视频上训练,不仅超越传统2D模型,将形状、结构及时间变化细节纳入预测,还实现了对具身智能体的精确逆动力学建模。首先,利用现成模型在现有机器人操作视频数据集上补充深度和法向信息;接着,在标注数据集上微调视频生成模型,联合预测每帧的RGB-DN;最后,提出算法将生成的RGB、深度和法向视频直接转换为高质量4D世界场景。该方法确保了具身场景中4D预测的时空一致性,支持具身环境的新视角合成,并使策略学习性能显著优于基于先前视频的世界模型。

原文摘要 · Abstract (English)

This paper presents an effective approach for learning novel 4D embodied world models, which predict the dynamic evolution of 3D scenes over time in response to an embodied agent's actions, providing both spatial and temporal consistency. We propose to learn a 4D world model by training on RGB-DN (RGB, Depth, and Normal) videos. This not only surpasses traditional 2D models by incorporating detailed shape, configuration, and temporal changes into their predictions, but also allows us to effectively learn accurate inverse dynamic models for an embodied agent. Specifically, we first extend existing robotic manipulation video datasets with depth and normal information leveraging off-the-shelf models. Next, we fine-tune a video generation model on this annotated dataset, which jointly predicts RGB-DN (RGB, Depth, and Normal) for each frame. We then present an algorithm to directly convert generated RGB, Depth, and Normal videos into a high-quality 4D scene of the world. Our method ensures temporal and spatial coherence in 4D scene predictions from embodied scenarios, enables novel view synthesis for embodied environments, and facilitates policy learning that significantly outperforms those derived from prior video-based world models.

4D建模具身智能世界模型视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。