通过人体-机器人视频对,实现细粒度动作迁移与零样本泛化。
Human2Robot: Learning Robot Actions from Paired Human-Robot Videos
- 将人机对齐视为条件视频生成问题,学习机器人动作隐式表征。
- 在2600个同步视频上训练,实现在新位置、新物体上的一次泛化。
- 适合机器人动作学习、远程操控及跨任务迁移研究者。
从人类示范中提炼知识是机器人学习与执行的有效途径。现有方法通常依赖粗略对齐的视频对,仅能学习全局或任务级特征,忽略复杂操作所需的细粒度帧级动态,难以泛化到新任务。我们认为这一局限源于数据集不足与方法相互制约的恶性循环。为此,我们提出范式转变:将细粒度人机对齐视为条件视频生成问题。首先引入H&R数据集,使用虚拟现实遥操作系统采集2600个精确同步的人体与机器人运动视频。随后提出Human2Robot框架,利用视频预测模型从人类输入生成机器人视频,从而学习丰富的隐式动作表征,并指导解耦的动作解码器。真实世界实验表明,该方法不仅在已见任务上表现优异,还能实现显著的一次泛化,适用于新位置、新物体、新实例乃至新任务类别。
原文摘要 · Abstract (English)
Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a result, they tend to neglect the fine-grained frame-level dynamics required for complex manipulation and generalization to novel tasks. We posit that this limitation stems from a vicious circle of inadequate datasets and the methods they inspire. To break this cycle, we propose a paradigm shift that treats fine-grained human-robot alignment as a conditional video generation problem. To this end, we first introduce H&R, a novel third-person dataset containing 2,600 episodes of precisely synchronized human and robot motions, collected using a VR teleoperation system. We then present Human2Robot, a framework designed to leverage this data. Human2Robot employs a Video Prediction Model to learn a rich and implicit representation of robot dynamics by generating robot videos from human input, which in turn guides a decoupled action decoder. Our real-world experiments demonstrate that this approach not only achieves high performance on seen tasks but also exhibits significant one-shot generalization to novel positions, objects, instances, and even new task categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。