arXiv:2606.19333cs.ROcs.CV2026-06被引 6

从日常人类视频中提取可执行的精细操作数据,让机器人直接模仿真人动作。

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos

论文配图:Do as I Do: Dexterous Manipulation Data from Everyday Human Videos
图 1 · 摘自论文原文
  • 通过重建人类视频中的手物交互,映射到机械手动作序列。
  • 在真实数据集上比现有方法更准确估计手物交互与操作轨迹。
  • 适合想用真实视频训练灵巧机器人的人工智能研究者。

如何规模化生成适用于灵巧多指机器手的操纵数据?近年来,从人类视频中学习成为可能路径。然而,手物交互估计困难及人机身体差异阻碍了单目RGB视频作为主要数据源的应用。本文提出DO AS I DO算法,从多种第一人称和第三人称户外视频中重建手物交互,并将其重定向为真实世界可执行的动作序列,从而从异构人类视频中生成完整的机器人操纵数据。实验表明,该方法在具有真实标注的数据集以及在线采集视频数据集上均优于现有最先进方法。研究还为实践者提供了采集人类操纵数据的有效策略指南。

原文摘要 · Abstract (English)

How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data. In this work, we present DO AS I DO, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. DO AS I DO reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos. Overall, DO AS I DO outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as we show in experiments on datasets with ground truths and on a dataset of video clips collected online. Our experiments enable us to propose an efficacy playbook for practitioners collecting human data for manipulation.

灵巧操作视频生成机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。