用大模型自动找关键点,让机器人学技能更省力、泛化更强。
Keypoint Abstraction using Large Models for Object-Relative Imitation Learning
- 用视觉语言大模型自动生成任务相关且跨实例一致的关键点
- 仅需少量示范数据即可在真实场景中实现跨物体、跨视角泛化
- 无需人工标注,适合复杂多变环境下的机器人模仿学习
在多样化任务与环境中实现对新物体构型和实例的泛化,是机器人领域的核心挑战。基于关键点的表征已被证明能有效捕捉物体核心特征,并建立动作预测的参考系,从而实现数据高效的学习机器人技能。然而,其依赖人工设计和额外人工标注的特性限制了可扩展性。本文提出KALM框架,利用预训练的视觉语言大模型(LM)自动生成任务相关的、跨实例一致的关键点。KALM通过大模型生成候选关键点,并结合少量机器人示范数据验证其鲁棒性与一致性。基于生成的关键点,可训练以关键点为中心的策略模型,实现对不同物体姿态、相机视角及相似功能形状物体的高效泛化。该方法在真实世界中表现优异,仅需少量示范即可适应多种任务与环境,且无需额外标签。
原文摘要 · Abstract (English)
Generalization to novel object configurations and instances across diverse tasks and environments is a critical challenge in robotics. Keypoint-based representations have been proven effective as a succinct representation for capturing essential object features, and for establishing a reference frame in action prediction, enabling data-efficient learning of robot skills. However, their manual design nature and reliance on additional human labels limit their scalability. In this paper, we propose KALM, a framework that leverages large pre-trained vision-language models (LMs) to automatically generate task-relevant and cross-instance consistent keypoints. KALM distills robust and consistent keypoints across views and objects by generating proposals using LMs and verifies them against a small set of robot demonstration data. Based on the generated keypoints, we can train keypoint-conditioned policy models that predict actions in keypoint-centric frames, enabling robots to generalize effectively across varying object poses, camera views, and object instances with similar functional shapes. Our method demonstrates strong performance in the real world, adapting to different tasks and environments from only a handful of demonstrations while requiring no additional labels. Website: https://kalm-il.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。