通过想象目标状态提升机器人抓取的精准度
Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation
- 用语言指令生成目标图像和3D点云作为先验知识
- 通过软姿态监督使动作与物体变化保持一致,减少误差
- 在仿真和真实场景中均优于现有方法,适合复杂抓取任务
关系性物体重排任务(如将花插入花瓶)要求机器人进行精确的语义与几何推理。现有方法或依赖预收集示范,难以捕捉复杂几何约束;或生成目标状态观测以获取语义与几何知识,但未显式关联物体变换与动作预测,导致生成噪声引发错误。为此,我们提出Imagine2Act,一种融合物体语义与几何约束的3D模仿学习框架。首先根据语言指令生成想象中的目标图像,并重建对应的3D点云,提供鲁棒的语义与几何先验。这些想象的目标点云作为策略模型的额外输入,同时采用带有软姿态监督的物体-动作一致性策略,显式对齐预测的末端执行器运动与生成的物体变换。该设计使Imagine2Act能够推理物体间的语义与几何关系,并在多样化任务中预测准确动作。仿真与真实世界实验表明,Imagine2Act优于以往最先进策略。更多可视化见 https://sites.google.com/view/imagine2act。
原文摘要 · Abstract (English)
Relational object rearrangement (ROR) tasks (e.g., insert flower to vase) require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches either rely on pre-collected demonstrations that struggle to capture complex geometric constraints or generate goal-state observations to capture semantic and geometric knowledge, but fail to explicitly couple object transformation with action prediction, resulting in errors due to generative noise. To address these limitations, we propose Imagine2Act, a 3D imitation-learning framework that incorporates semantic and geometric constraints of objects into policy learning to tackle high-precision manipulation tasks. We first generate imagined goal images conditioned on language instructions and reconstruct corresponding 3D point clouds to provide robust semantic and geometric priors. These imagined goal point clouds serve as additional inputs to the policy model, while an object-action consistency strategy with soft pose supervision explicitly aligns predicted end-effector motion with generated object transformation. This design enables Imagine2Act to reason about semantic and geometric relationships between objects and predict accurate actions across diverse tasks. Experiments in both simulation and the real world demonstrate that Imagine2Act outperforms previous state-of-the-art policies. More visualizations can be found at https://sites.google.com/view/imagine2act.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。