仅用一段人类视频,就能训练出能泛化的双手操作机器人策略。
Parse-Augment-Distill: Learning Generalizable Bimanual Visuomotor Policies from Single Human Video
- 从单段人类视频提取关键点动作轨迹,转为机器人可执行指令。
- 无需仿真,通过双臂任务规划大规模增强示范数据,成功率超现有方法。
- 适合快速部署双手操作任务,对场景和物体变化有强泛化能力。
从专家示范中学习视觉运动策略是现代机器人研究的重要方向,但现有方法多依赖大量遥控数据采集,且泛化能力差。虽然利用人类视频和示范增强技术可缓解数据瓶颈,但后者常需昂贵的仿真推演,引入模拟到现实的差距。与此同时,关键点等替代状态表示在类别级泛化方面表现优异。本文提出统一框架PAD(Parse-Augment-Distill),仅需一段人类视频即可学习可泛化的双手操作策略。该方法包含三步:(a) 将人类视频解析为机器人可执行的关键点-动作轨迹;(b) 采用双臂任务与运动规划,在无需仿真器的情况下大规模增强示范;(c) 将增强后的轨迹蒸馏为关键点条件策略。实验证明,PAD在成功率和样本/成本效率上均优于依赖仿真推演的图像策略方法。我们在六种真实世界双手任务中部署该框架,包括倒饮料、清垃圾、开容器等,生成的一次性策略可在未见过的空间布局、物体实例和背景干扰下成功泛化。
原文摘要 · Abstract (English)
Learning visuomotor policies from expert demonstrations is an important frontier in modern robotics research, however, most popular methods require copious efforts for collecting teleoperation data and struggle to generalize out-ofdistribution. Scaling data collection has been explored through leveraging human videos, as well as demonstration augmentation techniques. The latter approach typically requires expensive simulation rollouts and trains policies with synthetic image data, therefore introducing a sim-to-real gap. In parallel, alternative state representations such as keypoints have shown great promise for category-level generalization. In this work, we bring these avenues together in a unified framework: PAD (Parse-AugmentDistill), for learning generalizable bimanual policies from a single human video. Our method relies on three steps: (a) parsing a human video demo into a robot-executable keypoint-action trajectory, (b) employing bimanual task-and-motion-planning to augment the demonstration at scale without simulators, and (c) distilling the augmented trajectories into a keypoint-conditioned policy. Empirically, we showcase that PAD outperforms state-ofthe-art bimanual demonstration augmentation works relying on image policies with simulation rollouts, both in terms of success rate and sample/cost efficiency. We deploy our framework in six diverse real-world bimanual tasks such as pouring drinks, cleaning trash and opening containers, producing one-shot policies that generalize in unseen spatial arrangements, object instances and background distractors. Supplementary material can be found in the project webpage https://gtziafas.github.io/PAD_project/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。