用语义关键点让机器人少学几遍就能灵活完成复杂操作。
SKIL: Semantic Keypoint Imitation Learning for Generalizable Data-efficient Manipulation
- 通过视觉大模型自动提取语义关键点,构建高效模仿学习的描述符。
- 仅需30次演示即在挂毛巾任务中达70%成功率,性能是基线方法的两倍。
- 支持跨机器人平台迁移,连人类视频也能提升学习效果。
现实任务如衣物整理和桌面重排要求机器人具备泛化性、高精度与长时程动作能力。尽管模仿学习在教授新技能方面有效,但复杂任务仍需大量专家示范数据,导致样本复杂度高、数据收集成本昂贵。为此,我们提出语义关键点模仿学习(SKIL),借助视觉基础模型自动获取语义关键点,并构建其描述符,实现复杂机器人任务的高效模仿学习,显著降低样本需求。真实实验表明,SKIL在抓取杯子或鼠标等任务中性能提升一倍,对物体变化、环境扰动及干扰物具有极强鲁棒性。在以往方法完全失效的长时程任务如将毛巾挂上衣架中,仅需30次示范即达到70%平均成功率。此外,由于语义关键点的抽象特性,SKIL天然支持跨平台学习,即使使用人类视频也显著提升学习性能。所有结果证明SKIL在实现数据高效泛化机器人学习方面取得巨大成功。可视化与代码见:https://skil-robotics.github.io/SKIL-robotics/。
原文摘要 · Abstract (English)
Real-world tasks such as garment manipulation and table rearrangement demand robots to perform generalizable, highly precise, and long-horizon actions. Although imitation learning has proven to be an effective approach for teaching robots new skills, large amounts of expert demonstration data are still indispensible for these complex tasks, resulting in high sample complexity and costly data collection. To address this, we propose Semantic Keypoint Imitation Learning (SKIL), a framework which automatically obtains semantic keypoints with the help of vision foundation models, and forms the descriptor of semantic keypoints that enables efficient imitation learning of complex robotic tasks with significantly lower sample complexity. In real-world experiments, SKIL doubles the performance of baseline methods in tasks such as picking a cup or mouse, while demonstrating exceptional robustness to variations in objects, environmental changes, and distractors. For long-horizon tasks like hanging a towel on a rack where previous methods fail completely, SKIL achieves a mean success rate of 70\% with as few as 30 demonstrations. Furthermore, SKIL naturally supports cross-embodiment learning due to its semantic keypoints abstraction. Our experiments demonstrate that even human videos bring considerable improvement to the learning performance. All these results demonstrate the great success of SKIL in achieving data-efficient generalizable robotic learning. Visualizations and code are available at: https://skil-robotics.github.io/SKIL-robotics/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。