用关键点模仿学习提升机器人操控的泛化能力
On the Generalization Capabilities, Design Choices and Limitations of Keypoint Imitation Learning

- 基于视觉基础模型提取关键点,实现数据高效模仿学习
- 在5个任务中达成75%成功率,显著优于RGB基线(47%)
- 适合需要少样本泛化的机器人操控场景
基于RGB的模仿学习需大量演示才能泛化到未见物体或场景,促使研究者探索中间表示以提升机器人操控的泛化能力。视觉基础模型可实现单次演示提取关键点,提供此类表示。然而,如何最优集成这些模型以及它们何时优于其他表示仍不明确。本文综合先前关键点模仿学习(KIL)方法,探究多种设计选择,提供实用指导。基于超过2000次真实世界运行,评估KIL在未见物体和场景变化下的泛化能力。结果表明,KIL在五个任务上总体成功率达75%,显著优于RGB基线(47%),与S2-diffusion性能相当(73%)。同时,我们分析了用于关键点提取的基础模型的局限性,并将KIL扩展至多物体实例任务。结果确认KIL是数据高效的机器人学习方法,但未超越其他表示,且继承了基础模型的固有缺陷。所有演示视频、数据及结果详见https://kil-manipulation.github.io/。
原文摘要 · Abstract (English)
RGB-based imitation learning requires many demonstrations to generalize to unseen objects or scenes, motivating research into intermediate representations to improve generalization for robotic manipulation. Visual foundation models enable one-shot extraction of keypoints to provide such representation. However, it remains unclear how to integrate them into imitation learning optimally and when they outperform alternative representations. We combine approaches from previous works on keypoint imitation learning (KIL) and investigate several design choices to provide practical guidelines. Using over 2000 real-world rollouts, we also assess the generalization capabilities of KIL to unseen objects and scene variations. KIL achieves a 75% overall success rate across five tasks, significantly outperforming the RGB baseline (47%) and performing on par with S2-diffusion (73%). Finally, we explore the limitations of the foundation models used for keypoint extraction and extend KIL to tasks with multiple object instances. Our results confirm KIL as a data-efficient approach for robot learning, though it does not outperform alternative representations and inherits limitations of the foundation models used for keypoint extraction. All rollout videos, demonstrations, and results are available at https://kil-manipulation.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。