用真人视频训练通用灵巧手控制,降低数据成本并提升跨手泛化能力。
UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos
- 基于真人第一视角视频构建50K条轨迹数据集,实现人手到机器人手的精准动作迁移。
- 提出统一动作空间FAAS,使不同灵巧手间动作可迁移,任务成功率超81%。
- 配套便携采集系统,支持人机数据协同训练,适合机器人灵巧操作研究者。
由于真实机器人遥操作数据成本高、手部形态多样且控制维度高,灵巧操作仍具挑战。本文提出UniDex,一个包含大规模机器人中心数据集、统一视觉-语言-动作(VLA)策略及实用的人类数据采集方案的机器人基础套件。首先,构建UniDex-Dataset,涵盖8种灵巧手(6–24自由度)的50,000条轨迹,源自第一视角人类视频。通过人机协同重定向流程,对齐指尖轨迹并保留合理的手物接触;在显式3D点云上操作,并掩膜人类手部以缩小运动学与视觉差距。其次,提出功能-执行器对齐空间(FAAS),将功能相似的执行器映射至共享坐标,实现跨手迁移。基于FAAS参数化,训练了预训练于UniDex-Dataset的UniDex-VLA模型,并通过任务示范微调。此外,开发了UniDex-Cap,一种便携同步采集系统,记录RGB-D流与人类手姿,转换为机器人可执行轨迹,支持人-机数据联合训练,减少对昂贵机器人演示的依赖。在两种不同灵巧手上进行复杂工具使用任务测试,UniDex-VLA平均任务进度达81%,显著优于现有VLA基线,展现出强空间、物体及零样本跨手泛化能力。UniDex-Dataset、UniDex-VLA与UniDex-Cap共同构成可扩展的通用灵巧操作基础框架。
原文摘要 · Abstract (English)
Dexterous manipulation remains challenging due to the cost of collecting real-robot teleoperation data, the heterogeneity of hand embodiments, and the high dimensionality of control. We present UniDex, a robot foundation suite that couples a large-scale robot-centric dataset with a unified vision-language-action (VLA) policy and a practical human-data capture setup for universal dexterous hand control. First, we construct UniDex-Dataset, a robot-centric dataset over 50K trajectories across eight dexterous hands (6--24 DoFs), derived from egocentric human video datasets. To transform human data into robot-executable trajectories, we employ a human-in-the-loop retargeting procedure to align fingertip trajectories while preserving plausible hand-object contacts, and we operate on explicit 3D pointclouds with human hands masked to narrow kinematic and visual gaps. Second, we introduce the Function-Actuator-Aligned Space (FAAS), a unified action space that maps functionally similar actuators to shared coordinates, enabling cross-hand transfer. Leveraging FAAS as the action parameterization, we train UniDex-VLA, a 3D VLA policy pretrained on UniDex-Dataset and finetuned with task demonstrations. In addition, we build UniDex-Cap, a simple portable capture setup that records synchronized RGB-D streams and human hand poses and converts them into robot-executable trajectories to enable human-robot data co-training that reduces reliance on costly robot demonstrations. On challenging tool-use tasks across two different hands, UniDex-VLA achieves 81% average task progress and outperforms prior VLA baselines by a large margin, while exhibiting strong spatial, object, and zero-shot cross-hand generalization. Together, UniDex-Dataset, UniDex-VLA, and UniDex-Cap provide a scalable foundation suite for universal dexterous manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。