用人体动作桥接机器人操控,让双臂机械手更高效学习人类技巧。
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

- 提出相对腕部平移作为人机共用的动作表示,避免手姿噪声问题。
- 在新任务上,相比传统6自由度动作,技能迁移成功率提升显著。
- 适合研究人机协作、机器人模仿学习的学者与开发者参考。
我们研究如何将人类动作中的新操控技能迁移到具有并行夹持器的双臂机器人上。人类动作数据廉价、丰富且多样,是扩展机器人学习的重要资源。然而,将人类技能迁移到机器人仍具挑战:以往方法将人类视为另一种双臂6自由度体感,但手部姿态估计存在噪声,且人类手指接触模式与平行夹持器根本不同。因此,从人类数据中学习包含旋转的动作信号效果不佳。为此,我们提出一种桥接动作表示:以初始头戴相机坐标系下的相对腕部平移为动作空间,该空间同时适用于人类和机器人。为应对不同体感中某些动作成分缺失的问题,构建了类似$π_0$的视觉-语言-动作模型,采用交错动作标记与注意力掩码机制。在一系列新型双臂操控任务上,该方法比噪声较大的6自由度人类动作更有效地实现技能迁移,并随人类数据量增加而持续提升。
原文摘要 · Abstract (English)
We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Human action data is cheap, abundant, and diverse, making it one of the most promising resources for scaling up robot learning. Yet transferring skills from humans to robots remains hard: most prior work treats humans as just another bi-manual 6DoF embodiment, where hand-pose estimates are noisy and the contact patterns of human fingers differ fundamentally from those of a parallel gripper. We argue that learning rotation-inclusive action signals from human data is therefore sub-optimal, and instead propose a bridging action representation: the relative wrist translation within the initial head-camera frame, an action space shared by humans and robots. To handle the potential absence of certain action components in different embodiments, we build a $π_0$-like vision-language-action model with interleaved action tokens and attention masking. On a suite of novel bi-manual manipulation tasks, our bridging action transfers human manipulation knowledge to robots far more effectively than noisy 6DoF human actions and scales with the amount of human data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。