让机器人通过单视角预测完整4D场景并反推操作动作
MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation
- 用多视角融合与跨模态对齐,从单视角生成完整4D动态场景
- 在三个数据集上实现高质量4D生成和下游操控任务性能提升
- 适合做具身智能、机器人视觉-动作协同研究的学者参考
基于世界模型的“想象后执行”范式为机器人操作提供了新思路,但现有方法通常仅支持纯图像预测或部分3D几何推理,难以完整预测4D场景动态。本文提出一种新型具身4D世界模型,可实现几何一致、任意视角的RGBD生成:仅需单视角RGBD输入,模型即可推断其他视角,并通过反投影与融合构建更完整的时空3D结构。为高效学习多视角、跨模态生成,设计了显式的跨视角与跨模态特征融合机制,联合促进RGB与深度的一致性,并确保各视角间的几何对齐。在预测之外,传统逆动力学方法因解不唯一而效果不佳。为此,提出测试时动作优化策略:通过生成模型反向传播,推断匹配预测未来的轨迹级潜在变量;再由残差逆动力学模型将其转化为可执行的动作。在三个数据集上的实验表明,该方法在4D场景生成与下游操控任务中均表现优异,消融实验揭示了关键设计的有效性。
原文摘要 · Abstract (English)
World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely image-based forecasting or reasoning over partial 3D geometry, limiting their ability to predict complete 4D scene dynamics. This work proposes a novel embodied 4D world model that enables geometrically consistent, arbitrary-view RGBD generation: given only a single-view RGBD observation as input, the model imagines the remaining viewpoints, which can then be back-projected and fused to assemble a more complete 3D structure across time. To efficiently learn the multi-view, cross-modality generation, we explicitly design cross-view and cross-modality feature fusion that jointly encourage consistency between RGB and depth and enforce geometric alignment across views. Beyond prediction, converting generated futures into actions is often handled by inverse dynamics, which is ill-posed because multiple actions can explain the same transition. We address this with a test-time action optimization strategy that backpropagates through the generative model to infer a trajectory-level latent best matching the predicted future, and a residual inverse dynamics model that turns this trajectory prior into accurate executable actions. Experiments on three datasets demonstrate strong performance on both 4D scene generation and downstream manipulation, and ablations provide practical insights into the key design choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。