用扩散模型预测物体交互的潜在空间,提升机器人操控预测精度
LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation
- 在预训练视觉模型的潜在空间中建模未来状态演化
- 在真实场景中使策略性能提升20%,合成任务提升27.9%
- 适合需要高精度动作预测的机器人系统研究者
预测性操控近年来在具身智能领域受到广泛关注,因其可通过预测状态提升机器人策略性能。然而,从世界模型生成机器人-物体交互的准确未来视觉状态仍是难题,尤其在像素级表示上。为此,我们提出LaDi-WM,一种基于扩散模型的世界模型,用于预测未来状态的潜在空间。该模型利用与预训练视觉基础模型(VFMs)对齐的潜在空间,包含几何特征(基于DINO)和语义特征(基于CLIP)。我们发现,预测潜在空间的演变比直接预测像素级图像更易学习且更具泛化性。基于LaDi-WM,我们设计了一种扩散策略,通过引入预测状态迭代优化输出动作,从而生成更一致、准确的结果。在合成与真实世界基准上的大量实验表明,LaDi-WM在LIBERO-LONG基准上使策略性能提升27.9%,在真实场景中提升20%。此外,我们的世界模型与策略在真实实验中展现出出色的泛化能力。
原文摘要 · Abstract (English)
Predictive manipulation has recently gained considerable attention in the Embodied AI community due to its potential to improve robot policy performance by leveraging predicted states. However, generating accurate future visual states of robot-object interactions from world models remains a well-known challenge, particularly in achieving high-quality pixel-level representations. To this end, we propose LaDi-WM, a world model that predicts the latent space of future states using diffusion modeling. Specifically, LaDi-WM leverages the well-established latent space aligned with pre-trained Visual Foundation Models (VFMs), which comprises both geometric features (DINO-based) and semantic features (CLIP-based). We find that predicting the evolution of the latent space is easier to learn and more generalizable than directly predicting pixel-level images. Building on LaDi-WM, we design a diffusion policy that iteratively refines output actions by incorporating forecasted states, thereby generating more consistent and accurate results. Extensive experiments on both synthetic and real-world benchmarks demonstrate that LaDi-WM significantly enhances policy performance by 27.9\% on the LIBERO-LONG benchmark and 20\% on the real-world scenario. Furthermore, our world model and policies achieve impressive generalizability in real-world experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。