将单视角机器人视频生成多视角4D动态场景,提升机器人感知与决策能力。
Embody4D: A Generalist Data Engine for Embodied 4D World Modeling

- 通过3D感知合成管道构建多样化机器人场景数据集,增强泛化性。
- 引入隐空间置信度调制策略,保证生成视频时空一致性。
- 加入交互注意力机制,精准捕捉机械臂操作区域,提升动作真实性。
具身智能体需要鲁棒且全面的三维时空表征以支持空间推理、操作理解及下游决策。然而现有机器人数据通常来自固定或稀疏视角,仅提供部分且视角依赖的观测,限制了多视角感知和跨视角泛化。由于真实环境中难以采集更多视角,我们提出Embody4D,一种专用于具身场景的视频到视频世界模型,可将单目机器人视频转换为灵活目标视角的多视角视频。首先,为解决训练数据稀缺问题,设计3D感知组合合成流程,融合不同机器人手臂与多样背景,促进广泛泛化。其次,为保障几何稳定性,提出隐空间置信度感知的专家调制策略,估计扭曲隐式先验的可靠性,并自适应地将区域路由至复制、修复或补全专家,实现时空一致的4D生成。最后,为增强操作保真度,引入交互感知注意力机制,显式关注机器人交互区域。大量实验表明,Embody4D在视觉评估基准上达到顶尖性能,模拟与真实机器人实验进一步验证其作为高质量、视角一致视频数据引擎的有效性,可赋能下游机器人规划与学习。
原文摘要 · Abstract (English)
Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However, existing robot data are typically captured from fixed or sparse viewpoints, providing only partial and view-dependent observations, which limits multi-view perception and generalization across viewpoints. Given the difficulty of collecting additional viewpoints in real-world settings, we propose Embody4D, a dedicated video-to-video world model for embodied scenarios to bridge this observation gap by transforming a monocular robot video into novel-view videos from flexible target camera viewpoints. First, to tackle training data scarcity, we introduce a 3D-aware compositional synthesis pipeline to curate a heterogeneous dataset compositing cross-embodiment robotic arms with diverse backgrounds, promoting broad generalization. Second, to enforce geometric stability, we devise a latent confidence-aware expert modulation strategy, which estimates the reliability of warped latent priors and adaptively routes regions to copy, repair, or inpaint experts for spatiotemporally consistent 4D generation. Finally, to enhance the fidelity of the manipulation, we incorporate an interaction-aware attention mechanism that explicitly attends to the robotic interaction regions. Extensive experiments show that Embody4D achieves state-of-the-art performance on visual evaluation benchmarks, while both simulated and real-world robotic experiments further demonstrate its effectiveness as a robust data engine for synthesizing high-fidelity, view-consistent videos that empower downstream robotic planning and learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。