arXiv:2504.12609cs.ROcs.AI2025-04被引 50

仅用一段人类操作视频,让机器人学会复杂抓取与操作。

Crossing the Human-Robot Embodiment Gap with Sim-to-Real RL using One Human Demonstration

  • 从单段人类动作视频提取物体轨迹与手部初始位姿,构建仿真训练信号。
  • 在仿真中通过强化学习实现跨人机形态差距的策略训练,成功率超基线55%以上。
  • 无需穿戴设备或大量数据,适合快速部署新任务的机器人应用。

教会机器人灵巧操作通常需数百次穿戴设备或远程操控示范,难以规模化。视频形式的人机交互数据更易获取,但因缺乏明确动作标签及人机形态差异,直接用于机器人学习困难。本文提出 Human2Sim2Robot 框架,仅需一段 RGB-D 人类示范视频即可训练灵巧操作策略。方法从视频中提取:(1) 物体位姿轨迹以定义与形态无关的物体中心奖励;(2) 操作前手部姿态,用于引导强化学习中的探索。该机制使策略学习无需任务特定奖励调参。在单次人类示范条件下,该方法在抓取、非握持操作及多步骤任务上,分别优于对象感知回放超过55%,优于模仿学习超过68%。网站:https://human2sim2robot.github.io

原文摘要 · Abstract (English)

Teaching robots dexterous manipulation skills often requires collecting hundreds of demonstrations using wearables or teleoperation, a process that is challenging to scale. Videos of human-object interactions are easier to collect and scale, but leveraging them directly for robot learning is difficult due to the lack of explicit action labels and human-robot embodiment differences. We propose Human2Sim2Robot, a novel real-to-sim-to-real framework for training dexterous manipulation policies using only one RGB-D video of a human demonstrating a task. Our method utilizes reinforcement learning (RL) in simulation to cross the embodiment gap without relying on wearables, teleoperation, or large-scale data collection. From the video, we extract: (1) the object pose trajectory to define an object-centric, embodiment-agnostic reward, and (2) the pre-manipulation hand pose to initialize and guide exploration during RL training. These components enable effective policy learning without any task-specific reward tuning. In the single human demo regime, Human2Sim2Robot outperforms object-aware replay by over 55% and imitation learning by over 68% on grasping, non-prehensile manipulation, and multi-step tasks. Website: https://human2sim2robot.github.io

机器人学习强化学习单次示范仿真迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。