从普通视频学复杂抓取,无需传感器也能让机械手灵活操作物体。
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
- 用单目视频重建人手与物体的3D运动轨迹,直接转给机械手
- 仿真中抓取成功率70.25%,真实世界达62.86%,优于传统方法15.87%
- 无需额外机器人数据,一个视频就能生成多样训练样本
多指机械手操作与抓取因动作空间高维及大规模训练数据难获取而具挑战。现有方法多依赖穿戴设备或专用传感设备进行人类远程操控,限制了可扩展性。本文提出VIDEOMANIP,一种无需设备的框架,直接从RGB人类视频学习灵巧操作。借助计算机视觉进展,该框架通过估计人体手部姿态、物体网格,将重建的人体运动重定向至机械手,实现操作学习。为使重建数据适用于灵巧操作训练,引入以交互为中心的抓握建模和接触优化,并设计演示合成策略,从单一视频生成多样化训练轨迹,实现无需额外机器人示范的泛化策略学习。仿真中,所学抓取模型在Inspire Hand上对20种不同物体实现70.25%的成功率;真实世界中,基于RGB视频训练的策略在LEAP Hand上完成七项任务,平均成功率达62.86%,较基于重定向的方法提升15.87%。项目视频见videomanip.github.io。
原文摘要 · Abstract (English)
Multi-finger robotic hand manipulation and grasping are challenging due to the high-dimensional action space and the difficulty of acquiring large-scale training data. Existing approaches largely rely on human teleoperation with wearable devices or specialized sensing equipment to capture hand-object interactions, which limits scalability. In this work, we propose VIDEOMANIP, a device-free framework that learns dexterous manipulation directly from RGB human videos. Leveraging recent advances in computer vision, VIDEOMANIP reconstructs explicit 3D robot-object trajectories from monocular videos by estimating human hand poses, object meshes, and retargets the reconstructed human motions to robotic hands for manipulation learning. To make the reconstructed robot data suitable for dexterous manipulation training, we introduce hand-object contact optimization with interaction-centric grasp modeling, as well as a demonstration synthesis strategy that generates diverse training trajectories from a single video, enabling generalizable policy learning without additional robot demonstrations. In simulation, the learned grasping model achieves a 70.25% success rate across 20 diverse objects using the Inspire Hand. In the real world, manipulation policies trained from RGB videos achieve an average 62.86% success rate across seven tasks using the LEAP Hand, outperforming retargeting-based methods by 15.87%. Project videos are available at videomanip.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。