用人体动作视频直接训练机器人手,实现快速精准抓取。
Bridging the Human to Robot Dexterity Gap through Object-Oriented Rewards
- 基于物体轨迹构建奖励函数,跨形态差异学习
- 单视频输入+一小时交互,四指机械手完成任务
- 4倍性能提升,适合多指机器人实时学习
从人类视频中直接训练机器人是机器人学与计算机视觉的新兴方向。尽管双指夹持器已有显著进展,但将此类方法应用于多指机械手仍具挑战性,主要因人手与机器手形态差异导致策略难以迁移。本文提出HuDOR,一种通过人体视频在线微调策略的技术。该方法利用现成点跟踪器提取物体运动轨迹,构建面向物体的奖励函数,在形态与视觉差异下仍提供有效学习信号。仅需一个任务视频(如轻启音乐盒),四指Allegro机械手经一小时在线交互即可掌握任务。在四个任务上的实验表明,HuDOR相比基线提升4倍性能。代码与视频已公开于https://object-rewards.github.io。
原文摘要 · Abstract (English)
Training robots directly from human videos is an emerging area in robotics and computer vision. While there has been notable progress with two-fingered grippers, learning autonomous tasks for multi-fingered robot hands in this way remains challenging. A key reason for this difficulty is that a policy trained on human hands may not directly transfer to a robot hand due to morphology differences. In this work, we present HuDOR, a technique that enables online fine-tuning of policies by directly computing rewards from human videos. Importantly, this reward function is built using object-oriented trajectories derived from off-the-shelf point trackers, providing meaningful learning signals despite the morphology gap and visual differences between human and robot hands. Given a single video of a human solving a task, such as gently opening a music box, HuDOR enables our four-fingered Allegro hand to learn the task with just an hour of online interaction. Our experiments across four tasks show that HuDOR achieves a 4x improvement over baselines. Code and videos are available on our website, https://object-rewards.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。