无需动作标签数据,让机器人从视频中学会新任务
Generalist Robot Manipulation beyond Action Labeled Data
- 用3D动态点云自监督训练,无需动作标注即可学习
- 在真实与仿真环境中实现无动作标签的新任务学习
- 适合缺乏标注数据的通用机器人场景
近期通用机器人操作方法依赖预训练视觉语言模型和大规模带动作标注的机器人示范数据,在零样本条件下完成多样化任务。但高质量动作标注数据的获取仍是关键挑战。为此,我们提出一种新方法,可利用无动作标签的人类或机器人操作视频,提升开放词汇性能,并实现数据高效学习新任务。该方法在手部或夹爪位置提取密集动态3D点云,通过自研3D动态预测器进行自监督;再用小规模标注数据微调该预测器为动作预测器。实验表明,该方法不仅可从无标签示范视频中学习,优化下游通用机器人策略,还能在真实与仿真环境中实现无动作标签的新任务泛化能力。
原文摘要 · Abstract (English)
Recent advances in generalist robot manipulation leverage pre-trained Vision-Language Models (VLMs) and large-scale robot demonstrations to tackle diverse tasks in a zero-shot manner. A key challenge remains: scaling high-quality, action-labeled robot demonstration data, which existing methods rely on for robustness and generalization. To address this, we propose a method that benefits from videos without action labels - featuring humans and/or robots in action - enhancing open-vocabulary performance and enabling data-efficient learning of new tasks. Our method extracts dense, dynamic 3D point clouds at the hand or gripper location and uses a proposed 3D dynamics predictor for self-supervision. This predictor is then tuned to an action predictor using a smaller labeled dataset for action alignment. We show that our method not only learns from unlabeled human and robot demonstrations - improving downstream generalist robot policies - but also enables robots to learn new tasks without action labels (i.e., out-of-action generalization) in both real-world and simulated settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。