arXiv:2410.20018cs.ROcs.AI2024-10ICRA被引 16

用过滤机制提升生成目标与机器人控制的匹配度,让智能体更稳地完成任务。

GHIL-Glue: Hierarchical Control with Filtered Subgoal Images

  • 通过筛选无效子目标,减少生成图像对低层策略的干扰。
  • 在模拟和真实环境中使任务成功率提升25%,达新基准水平。
  • 适合需要语言指令、零样本泛化的机器人操控场景。

在互联网规模数据上预训练的图像与视频生成模型能显著增强机器人学习系统的泛化能力。这些模型可作为高层规划器,为底层条件化策略生成中间子目标。然而,生成模型与底层控制器之间的接口常成为性能瓶颈:生成模型可能输出物理上不可行的逼真帧,混淆底层策略;而底层策略也可能对生成目标中的细微视觉瑕疵敏感。本文提出生成式分层模仿学习-粘合(GHIL-Glue),通过过滤无法推进任务的子目标,并增强底层策略对有害视觉伪影的鲁棒性,有效“粘合”语言条件化的图像/视频预测模型与底层目标条件化策略。在仿真与真实环境的大量实验中,GHIL-Glue在多个使用生成子目标的分层模型上实现25%的性能提升,在仅用单个RGB相机观测的CALVIN仿真基准上达到新最优水平。在物理实验中,其在4项语言指令操控任务中,有3项超越其他通用机器人策略,展现优异的零样本泛化能力。

原文摘要 · Abstract (English)

Image and video generative models that are pre-trained on Internet-scale data can greatly increase the generalization capacity of robot learning systems. These models can function as high-level planners, generating intermediate subgoals for low-level goal-conditioned policies to reach. However, the performance of these systems can be greatly bottlenecked by the interface between generative models and low-level controllers. For example, generative models may predict photorealistic yet physically infeasible frames that confuse low-level policies. Low-level policies may also be sensitive to subtle visual artifacts in generated goal images. This paper addresses these two facets of generalization, providing an interface to effectively "glue together" language-conditioned image or video prediction models with low-level goal-conditioned policies. Our method, Generative Hierarchical Imitation Learning-Glue (GHIL-Glue), filters out subgoals that do not lead to task progress and improves the robustness of goal-conditioned policies to generated subgoals with harmful visual artifacts. We find in extensive experiments in both simulated and real environments that GHIL-Glue achieves a 25% improvement across several hierarchical models that leverage generative subgoals, achieving a new state-of-the-art on the CALVIN simulation benchmark for policies using observations from a single RGB camera. GHIL-Glue also outperforms other generalist robot policies across 3/4 language-conditioned manipulation tasks testing zero-shot generalization in physical experiments.

机器人学习分层控制生成模型零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。