从真人视频直接学习双臂灵巧操作,无需传感器或标注
DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated Videos
- 直接用第三人称视频训练机器人双臂操作技能
- 在TACO数据集上姿态估计提升0.08(ADD-S)和0.12(VSD)
- 支持真实与合成视频,可生成大规模多样化训练数据
我们提出DexMan,一个自动化框架,将人类视觉示范转化为类人机器人在仿真环境中的双臂灵巧操作技能。该框架直接处理人类操纵刚性物体的第三人称视频,无需相机标定、深度传感器、扫描3D物体资产或真实手部与物体运动标注。不同于以往仅考虑简化浮动手部的方法,DexMan直接控制类人机器人,并引入基于接触的奖励机制,提升从野外视频中估算的噪声手物姿态中学习策略的性能。在TACO基准测试中,其对象姿态估计达到领先水平,绝对提升分别为0.08(ADD-S)和0.12(VSD)。同时,其强化学习策略在OakInk-v2任务上的成功率较之前方法提升19%。此外,DexMan可利用真实与合成视频生成技能,无需人工数据采集或昂贵动作捕捉,实现通用灵巧操作的大规模多样数据集构建。
原文摘要 · Abstract (English)
We present DexMan, an automated framework that converts human visual demonstrations into bimanual dexterous manipulation skills for humanoid robots in simulation. Operating directly on third-person videos of humans manipulating rigid objects, DexMan eliminates the need for camera calibration, depth sensors, scanned 3D object assets, or ground-truth hand and object motion annotations. Unlike prior approaches that consider only simplified floating hands, it directly controls a humanoid robot and leverages novel contact-based rewards to improve policy learning from noisy hand-object poses estimated from in-the-wild videos. DexMan achieves state-of-the-art performance in object pose estimation on the TACO benchmark, with absolute gains of 0.08 and 0.12 in ADD-S and VSD. Meanwhile, its reinforcement learning policy surpasses previous methods by 19% in success rate on OakInk-v2. Furthermore, DexMan can generate skills from both real and synthetic videos, without the need for manual data collection and costly motion capture, and enabling the creation of large-scale, diverse datasets for training generalist dexterous manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。