仅用一段视频教学,让双臂机器人学会复杂操作。
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
- 从单段双视角视频中提取双手动作模式并注入机器人
- 实现5个复杂长时序任务的高精度模仿,泛化能力强
- 适合想快速部署双臂机器人的研究者与工程师
双臂机器人操作因需协调双臂时空动作且动作空间维度高,长期面临挑战。以往方法依赖预定义动作分类或远程操控,往往缺乏简洁性、通用性与可扩展性。本文提出YOTO(You Only Teach Once),通过仅需一段双目手部运动视频,即可提取并注入双臂动作模式,教会双机器人完成多种复杂任务。基于关键帧运动轨迹,我们设计了一种快速生成多样化物体位置与形态训练样本的方法,用于在不同场景中学习定制化的双臂扩散策略(BiDP)。实验表明,YOTO能成功模仿5个复杂长时序双臂任务,在不同视觉和空间条件下具备强泛化能力,且在准确率与效率上优于现有视觉-动作模仿学习方法。
原文摘要 · Abstract (English)
Bimanual robotic manipulation is a long-standing challenge of embodied intelligence due to its characteristics of dual-arm spatial-temporal coordination and high-dimensional action spaces. Previous studies rely on pre-defined action taxonomies or direct teleoperation to alleviate or circumvent these issues, often making them lack simplicity, versatility and scalability. Differently, we believe that the most effective and efficient way for teaching bimanual manipulation is learning from human demonstrated videos, where rich features such as spatial-temporal positions, dynamic postures, interaction states and dexterous transitions are available almost for free. In this work, we propose the YOTO (You Only Teach Once), which can extract and then inject patterns of bimanual actions from as few as a single binocular observation of hand movements, and teach dual robot arms various complex tasks. Furthermore, based on keyframes-based motion trajectories, we devise a subtle solution for rapidly generating training demonstrations with diverse variations of manipulated objects and their locations. These data can then be used to learn a customized bimanual diffusion policy (BiDP) across diverse scenes. In experiments, YOTO achieves impressive performance in mimicking 5 intricate long-horizon bimanual tasks, possesses strong generalization under different visual and spatial conditions, and outperforms existing visuomotor imitation learning methods in accuracy and efficiency. Our project link is https://hnuzhy.github.io/projects/YOTO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。