用单视角人手视频生成机器人可执行的抓取数据,解决数据稀缺问题。
DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos
- 通过四阶段流程将人手视频转为物理合理的机器人动作数据
- 支持工具使用、长序列任务和精细操作等多样场景
- 无需额外标注,直接用于真实机器人零样本部署
数据稀缺严重制约双臂灵巧操作的泛化能力,因真实世界灵巧手数据采集成本高且耗时。人类操作视频作为操作知识的直接载体,具备大规模扩展机器人学习的潜力。然而,人类手与机器人灵巧手间存在显著形态差异,导致直接从人手视频预训练极为困难。为此,我们提出DexImit,一种自动化框架,仅需单视角人手操作视频,即可生成物理上合理且可执行的机器人数据,无需任何额外信息。该框架包含四个阶段:(1) 从任意视角重建近度量尺度的手-物交互;(2) 实现子任务分解与双臂调度;(3) 合成与示范一致的机器人轨迹;(4) 全面数据增强以实现零样本真实部署。基于此设计,DexImit可基于互联网视频或视频生成模型生成大规模机器人数据,适用于多种操作任务,包括工具使用(如切苹果)、长时序任务(如制作饮品)及精细操作(如叠杯)。
原文摘要 · Abstract (English)
Data scarcity fundamentally limits the generalization of bimanual dexterous manipulation, as real-world data collection for dexterous hands is expensive and labor-intensive. Human manipulation videos, as a direct carrier of manipulation knowledge, offer significant potential for scaling up robot learning. However, the substantial embodiment gap between human hands and robotic dexterous hands makes direct pretraining from human videos extremely challenging. To bridge this gap and unleash the potential of large-scale human manipulation video data, we propose DexImit, an automated framework that converts monocular human manipulation videos into physically plausible robot data, without any additional information. DexImit employs a four-stage generation pipeline: (1) reconstructing hand-object interactions from arbitrary viewpoints with near-metric scale; (2) performing subtask decomposition and bimanual scheduling; (3) synthesizing robot trajectories consistent with the demonstrated interactions; (4) comprehensive data augmentation for zero-shot real-world deployment. Building on these designs, DexImit can generate large-scale robot data based on human videos, either from the Internet or video generation models. DexImit is capable of handling diverse manipulation tasks, including tool use (e.g., cutting an apple), long-horizon tasks (e.g., making a beverage), and fine-grained manipulations (e.g., stacking cups).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。