arXiv:2409.12005cs.ROcs.AI2024-09被引 2

改进世界模型对位置信息的表示,提升机器人抓取任务表现

Representing Positional Information in Generative World Models for Object Manipulation

  • 用位置或隐变量条件化策略,显式建模物体位置信息
  • 在多个环境中优于现有模型基控制方法
  • 支持坐标或视觉目标的多模态目标设定,适合机器人操控研究

物体操作能力是具身智能体与世界交互的核心技能,尤其在机器人领域至关重要。准确预测物体交互结果对完成任务尤为关键。尽管基于模型的控制方法已用于操作任务,但其在精确操控方面仍存在不足。我们分析发现,问题根源在于当前世界模型对关键位置信息(尤其是目标定位信息)的表征不够充分。为此,提出一种通用方法,增强基于世界模型的智能体解决物体定位任务的能力。针对生成式世界模型,设计两种实现方式:位置条件化策略(PCP)和隐变量条件化策略(LCP)。其中,LCP利用以物体为中心的隐变量表示,显式捕捉物体位置信息,从而自然涌现出多模态能力,可支持通过空间坐标或视觉图像指定目标。在多个操作环境中的严格评估表明,该方法相较现有模型基控制方法表现更优。

原文摘要 · Abstract (English)

Object manipulation capabilities are essential skills that set apart embodied agents engaging with the world, especially in the realm of robotics. The ability to predict outcomes of interactions with objects is paramount in this setting. While model-based control methods have started to be employed for tackling manipulation tasks, they have faced challenges in accurately manipulating objects. As we analyze the causes of this limitation, we identify the cause of underperformance in the way current world models represent crucial positional information, especially about the target's goal specification for object positioning tasks. We introduce a general approach that empowers world model-based agents to effectively solve object-positioning tasks. We propose two declinations of this approach for generative world models: position-conditioned (PCP) and latent-conditioned (LCP) policy learning. In particular, LCP employs object-centric latent representations that explicitly capture object positional information for goal specification. This naturally leads to the emergence of multimodal capabilities, enabling the specification of goals through spatial coordinates or a visual goal. Our methods are rigorously evaluated across several manipulation environments, showing favorable performance compared to current model-based control approaches.

世界模型机器人操控位置表示多模态目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。