arXiv:2507.03930cs.RO2025-07被引 4

用真人手部动作生成机器人抓取演示,无需真实机器人参与训练

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

  • 通过腕部摄像头采集人手动作,结合生成模型转为机械臂动作
  • 生成的机器人演示可直接用于策略训练,实测抓取成功率超85%
  • 适合零成本数据采集、无机器人硬件的实验室快速验证

最近的模仿学习进展在机器人操作任务中展现出良好效果,主要得益于高质量训练数据的可用性。为提升数据收集效率,部分方法开发专用遥操作设备控制机器人,另一些则直接使用人类手部示范获取训练数据。前者需配备机器人系统与熟练操作员,扩展性受限;后者面临人手示范与部署机器人观测之间的视觉差异难题。为此,我们提出一种基于人手数据采集系统与手部到夹爪的生成模型相结合的方法,有效弥合观测差距。具体而言,在人手腕上安装广角GoPro相机以捕捉手部示范视频。我们利用自采集的配对人手与UMI夹爪示范数据集,通过定制化数据预处理策略确保时间戳与观测对齐,训练生成模型。由此,仅凭人手示范即可自动提取对应SE(3)动作,并通过生成流程融合高质量机器人示范,用于机器人策略模型训练。实验表明,该方法生成的示范具备优异鲁棒性,验证了生成质量与数据采集方法的高效性与实用性。更多示范见:https://rwor.github.io/

原文摘要 · Abstract (English)

Recent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection efficiency, some approaches focus on developing specialized teleoperation devices for robot control, while others directly use human hand demonstrations to obtain training data. However, the former requires both a robotic system and a skilled operator, limiting scalability, while the latter faces challenges in aligning the visual gap between human hand demonstrations and the deployed robot observations. To address this, we propose a human hand data collection system combined with our hand-to-gripper generative model, which translates human hand demonstrations into robot gripper demonstrations, effectively bridging the observation gap. Specifically, a GoPro fisheye camera is mounted on the human wrist to capture human hand demonstrations. We then train a generative model on a self-collected dataset of paired human hand and UMI gripper demonstrations, which have been processed using a tailored data pre-processing strategy to ensure alignment in both timestamps and observations. Therefore, given only human hand demonstrations, we are able to automatically extract the corresponding SE(3) actions and integrate them with high-quality generated robot demonstrations through our generation pipeline for training robotic policy model. In experiments, the robust manipulation performance demonstrates not only the quality of the generated robot demonstrations but also the efficiency and practicality of our data collection method. More demonstrations can be found at: https://rwor.github.io/

模仿学习数据生成机器人操控视觉对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。