用分层目标策略让机器人学会泛化操作,仅需少量人类示范即可适应新物体。
GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

- 高层策略预测3D操作目标,低层控制器执行具体动作,实现分层控制
- 在多个任务中表现优于平铺式扩散策略,鲁棒性显著提升
- 可直接利用人类视频示范,无需复杂动作重映射,适合快速迁移
我们提出GHOST框架,用于学习能超越训练分布的视觉-运动操作策略。该框架将控制分解为:(i) 高层策略,从多视角RGB-D观测中预测下一子目标,以3D末端执行器位姿的概率分布表示;(ii) 低层目标条件控制器,执行与具体机器人形态相关的动作。为使基于图像的策略适配3D目标,我们引入一个简单空间接口,将预测的目标投影到图像平面,以末端执行器热图形式表示。在一系列操作任务中,该分层结构相比平铺式扩散策略表现更优且更具鲁棒性。此外,该分层接口便于直接融入人类示范,无需依赖(噪声)动作重映射。由于子目标具有高度形态无关性,我们使用人类视频训练高层策略以定义技能的应用与组合方式,而保持低层策略仅用机器人数据训练。这种层级结构使得仅通过少量人类示范即可适应新物体和任务变化。
原文摘要 · Abstract (English)
We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control into (i) a high-level policy that predicts the next sub-goal as a distribution over 3D end-effector poses from multi-view RGB-D observations, and (ii) a low-level goal-conditioned controller that executes embodiment-specific actions. To condition image-based policies on 3D goals, we introduce a simple spatial interface that projects predicted goals into the image plane and represents them as end-effector heatmaps. Across a suite of manipulation tasks, this hierarchical factorization consistently improves performance and robustness compared to a flat Diffusion Policy. Further, we show that this hierarchical interface also makes it easy to incorporate human demonstrations without relying on (noisy) action retargeting. As sub-goals are largely embodiment-agnostic, we train the high-level policy on human video to specify how learned skills should be applied and composed, while keeping the low-level policy trained purely on robot data. This hierarchy enables adaptation to novel objects and task variations using a small number of human demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。