arXiv:2505.20795cs.RO2025-05中稿 · ICRA被引 4

看视频学任务:机器人无需新数据就能直接模仿人类示范。

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

  • 用跨视频预测学习人与机器人的联合表征
  • 在真实操作任务中实现零样本泛化,性能优于基线方法
  • 适合需要快速部署新任务的机器人应用

当前机器人学习方法通常依赖大量通过遥操作收集的机器人数据。面对新任务时,这类方法需重新采集遥操作数据并微调策略,而该过程繁琐且成本高。相比之下,人类仅需观看他人操作即可快速掌握新任务。本文提出一种两阶段框架,利用人类示范视频作为提示,训练具备泛化能力的机器人策略。第一阶段通过交叉预测训练视频生成模型,捕捉人类与机器人示范视频的联合表征;第二阶段采用新型原型对比损失,将学习到的表征与人机共享动作空间融合。在真实世界灵巧操作任务上的实证评估表明,所提方法具有出色的泛化能力。

原文摘要 · Abstract (English)

Recent robot learning methods commonly rely on imitation learning from massive robotic dataset collected with teleoperation. When facing a new task, such methods generally require collecting a set of new teleoperation data and finetuning the policy. Furthermore, the teleoperation data collection pipeline is also tedious and expensive. Instead, human is able to efficiently learn new tasks by just watching others do. In this paper, we introduce a novel two-stage framework that utilizes human demonstrations to learn a generalizable robot policy. Such policy can directly take human demonstration video as a prompt and perform new tasks without any new teleoperation data and model finetuning at all. In the first stage, we train video generation model that captures a joint representation for both the human and robot demonstration video data using cross-prediction. In the second stage, we fuse the learned representation with a shared action space between human and robot using a novel prototypical contrastive loss. Empirical evaluations on real-world dexterous manipulation tasks show the effectiveness and generalization capabilities of our proposed method.

机器人学习模仿学习视频提示零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。