arXiv:2509.10952cs.RO2025-09被引 27

用人类视频教机器人模仿动作,跨域适配更顺畅。

ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation

  • 通过动态时间规整映射人手动作到机器人关节
  • 混合插值生成中间域数据,提升成功率与执行平滑度
  • 无需大量机器人数据,适合多类型机械臂迁移

从大量人类视频中学习机器人操作,是替代昂贵机器人数据采集的可扩展方案。然而,视觉、形态和物理层面的领域差异阻碍了直接模仿。为此,我们提出 ImMimic,一种不依赖具体身体结构的联合训练框架,结合人类视频与少量遥操作机器人示范。ImMimic 使用动态时间规整(DTW)进行基于动作或视觉的映射,将重定向的人类手部姿态映射至机器人关节,并对成对的人类与机器人轨迹进行 MixUp 插值。关键洞察为:(1) 重定向的人类手部轨迹提供有效动作标签;(2) 映射数据上的插值生成中间域,促进联合训练中的平稳领域适应。在四个真实世界操作任务(拾取放置、推、锤击、翻转)及四种机器人本体(Robotiq、Fin Ray、Allegro、Ability)上评估显示,ImMimic 提升了任务成功率与执行平滑性,验证其在跨域适配方面的有效性。

原文摘要 · Abstract (English)

Learning robot manipulation from abundant human videos offers a scalable alternative to costly robot-specific data collection. However, domain gaps across visual, morphological, and physical aspects hinder direct imitation. To effectively bridge the domain gap, we propose ImMimic, an embodiment-agnostic co-training framework that leverages both human videos and a small amount of teleoperated robot demonstrations. ImMimic uses Dynamic Time Warping (DTW) with either action- or visual-based mapping to map retargeted human hand poses to robot joints, followed by MixUp interpolation between paired human and robot trajectories. Our key insights are (1) retargeted human hand trajectories provide informative action labels, and (2) interpolation over the mapped data creates intermediate domains that facilitate smooth domain adaptation during co-training. Evaluations on four real-world manipulation tasks (Pick and Place, Push, Hammer, Flip) across four robotic embodiments (Robotiq, Fin Ray, Allegro, Ability) show that ImMimic improves task success rates and execution smoothness, highlighting its efficacy to bridge the domain gap for robust robot manipulation. The project website can be found at https://sites.google.com/view/immimic.

机器人模仿跨域适配动作迁移视频学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。