用4D结构化潜在空间预测3D场景演化,提升机器人长程规划能力。
Structured 4D Latent Predictive Model for Robot Planning

- 构建4D结构化潜在空间,融合观测与文本指令预测3D场景演化
- 生成未来画面视觉质量高,3D一致性与多视角连贯性显著优于现有方法
- 适用于复杂操作任务,支持真实机器人平台部署,泛化能力强
视频预测模型正成为机器人领域的重要范式,为任务泛化、长时序规划和灵活决策提供新路径。然而,现有方法多基于2D视频序列,缺乏必要的3D几何理解,难以实现精准的空间推理与物理一致性。本文提出一种结构化4D潜在预测模型,通过条件化观测与文本指令,在结构化潜在空间中预测场景的3D结构演化。该表示对场景进行整体建模,可解码为多种3D格式,实现更完整且一致的3D场景理解。该模型作为规划器,生成的未来场景由目标条件逆动力学模块转化为可执行动作。实验表明,相比最先进视频规划方法,本模型生成的未来在视觉质量、3D一致性及多视角连贯性上均有显著提升。最终规划流水线在复杂操作任务中表现优异,对新视觉条件具备强泛化能力,并在真实机器人平台上验证有效。项目主页:https://structured-4d-model.github.io/
原文摘要 · Abstract (English)
Video predictive models are emerging as a powerful paradigm in robotics, offering a promising path toward task generalization, long-horizon planning, and flexible decision-making. However, prevailing approaches often operate on 2D video sequences, inherently lacking the 3D geometric understanding necessary for precise spatial reasoning and physical consistency. We introduce a Structured 4D Latent Predictive Model, which predicts the evolution of a scene's 3D structure in a structured latent space conditioned on observations and textual instructions. Our representation encodes the scene holistically and can be decoded into diverse 3D formats, enabling a more complete and 3D consistent scene understanding. This structured 4D latent predictive model serves as a planner, generating future scenes that are translated into executable actions by a goal-conditioned inverse dynamics module. Experiments demonstrate that our model generates futures with strong visual quality, substantially better 3D consistency and multi-view coherence compared to state-of-the-art video-based planners. Consequently, our full planning pipeline achieves superior performance on complex manipulation tasks, exhibits robust generalization to novel visual conditions, and proves effective on real-world robotic platforms. Our website is available at https://structured-4d-model.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。