用物体为中心的表示预测未来,让机器人更高效地听懂指令操作物体。
Object-Centric World Model for Language-Guided Manipulation
- 基于槽注意力构建物体中心表征,用语言指令引导未来状态预测。
- 在视觉-语言-动作控制任务中,样本和计算效率优于生成式模型。
- 适合需要精准识别物体的机械臂操作任务,泛化能力强。
世界模型对自动驾驶和机器人等领域的未来预测与规划至关重要。近年来,视频生成因扩散模型的成功而备受关注,但其计算开销大。为此,本文提出一种基于槽注意力的物体中心世界模型,将当前状态建模为物体中心表示,并在自然语言指令指导下预测未来状态。该方法相比生成式模型更紧凑、计算更高效,且能灵活响应语言指令,在物体识别关键的操控任务中表现优异。实验表明,该潜在预测世界模型在视觉-语言-运动控制任务中超越生成式世界模型,实现更高的样本与计算效率。我们还研究了该方法的泛化性能,并探索了利用物体中心表示预测动作的多种策略。
原文摘要 · Abstract (English)
A world model is essential for an agent to predict the future and plan in domains such as autonomous driving and robotics. To achieve this, recent advancements have focused on video generation, which has gained significant attention due to the impressive success of diffusion models. However, these models require substantial computational resources. To address these challenges, we propose a world model leveraging object-centric representation space using slot attention, guided by language instructions. Our model perceives the current state as an object-centric representation and predicts future states in this representation space conditioned on natural language instructions. This approach results in a more compact and computationally efficient model compared to diffusion-based generative alternatives. Furthermore, it flexibly predicts future states based on language instructions, and offers a significant advantage in manipulation tasks where object recognition is crucial. In this paper, we demonstrate that our latent predictive world model surpasses generative world models in visuo-linguo-motor control tasks, achieving superior sample and computation efficiency. We also investigate the generalization performance of the proposed method and explore various strategies for predicting actions using object-centric representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。