arXiv:2609.02531cs.CVcs.RO2026-09

用3D深度信息增强机器人动作预测模型,提升真实场景适应力。

Spatially Aware World Action Model via Geometric Latent Diffusion

论文配图:Spatially Aware World Action Model via Geometric Latent Diffusion
图 1 · 摘自论文原文
  • 将深度图融入预训练视频扩散模型,实现RGB与深度联合预测
  • 在RoboCasa和LIBERO-Plus上达最新最佳性能,真实机器人测试胜出
  • 无需微调编码器即可融合几何信息,保留原始视觉先验

世界动作模型(WAMs)利用大规模预训练视频扩散模型,联合预测未来观测与动作,继承了互联网级视频中的丰富视觉与物理先验,是机器人策略学习的有前景范式。然而现有模型仅基于RGB观测,未利用3D信息。为此,我们提出空间感知的世界动作模型(SA-WAM),将预训练视频模型改造为同时预测动作、RGB与深度图像,实现单个扩散主干中的3D感知建模与动作预测。通过非线性编码将无界深度信号映射到预训练VAE分词器期望的有界输入域,无需3D专属微调即可复用分词器,融入几何信息而不损失预训练先验。SA-WAM在RoboCasa与LIBERO-Plus基准上达到最优表现,同时提升未来状态预测能力。此外,在使用UR5机械臂的真实世界评估中,显著优于强基线,尤其在随机化环境中优势明显。我们分析了世界模型预测质量与轨迹成功间的相关性,为提升WAM性能提供洞见。

原文摘要 · Abstract (English)

World Action Models (WAMs) leverage the capabilities of large-scale pretrained video diffusion models to jointly predict future observations and actions, inheriting rich visual and physical priors from internet-scale video. This has made them a promising paradigm for robot policy learning, yet the prevailing models operate exclusively on RGB observations and do not leverage 3D information. To bridge this gap, we introduce a Spatially Aware World Action Model (SA-WAM), which repurposes a pretrained video model for joint action, RGB, and depth prediction, enabling 3D-aware world modeling and action prediction within a single diffusion backbone. We use a nonlinear encoding that maps the unbounded depth signal into the bounded input domain expected by the frozen VAE tokenizer. This allows us to reuse the tokenizer without 3D-specific fine-tuning, incorporating geometric information without sacrificing the pretrained priors. SA-WAM achieves state-of-the-art results on the RoboCasa and LIBERO-Plus benchmarks, while simultaneously improving future-state predictions. Furthermore, SA-WAM outperforms strong baselines in real-world evaluation using a UR5 robotic arm, with strong gains in randomized environments. We analyze the correlation between world model prediction quality and rollout success, providing insights into WAM performance and avenues for its improvement.

机器人学习扩散模型3D感知动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。