通过模仿人类动作关键点,让机器人在少样本下更准确地理解语言指令完成操作。
AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
- 用人类动作视频显式学习手部关键点,构建视觉语言模型
- 在CALVIN基准和真实场景中表现领先,少样本下优势显著
- 适合数据稀缺的机器人操控任务,尤其适用于少样本学习
视觉机器人操作(VRM)旨在使机器人根据机器人状态和视觉观察执行自然语言指令,但需要大量多模态数据。现有方法依赖大规模网络数据或隐式训练(如像素级未来帧预测),在机器人数据不足时泛化能力差。本文提出一种显式模仿人类动作的方法——基于类比推理的视觉机器人操作(AR-VRM)。通过从大规模人类动作视频中提取手部关键点,训练视觉语言模型以显式学习人类动作知识,并直接预测手部关键点。微调阶段,通过检索相似任务与历史观察的人类动作视频,建立人类手部关键点与机器人部件之间的类比推理映射。相比关注无关视觉线索的方法,本方法在CALVIN基准和真实世界实验中均取得领先性能。在少样本场景下,显著优于先前方法,验证了在数据稀缺条件下显式模仿人类动作的有效性。
原文摘要 · Abstract (English)
Visual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multi-modal data. To compensate for the deficiency of robot data, existing approaches have employed vision-language pretraining with large-scale data. However, they either utilize web data that differs from robotic tasks, or train the model in an implicit way (e.g., predicting future frames at the pixel level), thus showing limited generalization ability under insufficient robot data. In this paper, we propose to learn from large-scale human action video datasets in an explicit way (i.e., imitating human actions from hand keypoints), introducing Visual Robot Manipulation with Analogical Reasoning (AR-VRM). To acquire action knowledge explicitly from human action videos, we propose a keypoint Vision-Language Model (VLM) pretraining scheme, enabling the VLM to learn human action knowledge and directly predict human hand keypoints. During fine-tuning on robot data, to facilitate the robotic arm in imitating the action patterns of human motions, we first retrieve human action videos that perform similar manipulation tasks and have similar historical observations , and then learn the Analogical Reasoning (AR) map between human hand keypoints and robot components. Taking advantage of focusing on action keypoints instead of irrelevant visual cues, our method achieves leading performance on the CALVIN benchmark {and real-world experiments}. In few-shot scenarios, our AR-VRM outperforms previous methods by large margins , underscoring the effectiveness of explicitly imitating human actions under data scarcity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。