统一异构灵巧手状态表示,实现无重定向的高精度建模
UniDexTok: A Unified Dexterous Hand Tokenizer from Real Data

- 基于统一手部模型构建共享语义接口,标准化不同硬件手的状态
- 相比基线降低99%以上误差,重建精度达亚毫米级
- 支持跨型号手部数据迁移,适合新手型快速部署
灵巧手在精细操作中至关重要,但其硬件设计差异大,运动学结构、关节定义和自由度各不相同,难以建立统一的状态表示,导致数据分散且难用于联合训练。本文提出统一灵巧手模型(UDHM),将人手与机器人手状态映射至22自由度的共享语义接口。基于此,提出UniDexTok,一种无需重定向的实时状态分词器,直接从标准化真实关节数据学习具身体现条件的离散令牌。该方法无需仿真或重定向数据,即可实现异构灵巧手的统一表征。相比近期基线方法UniHM,UniDexTok将平均关节角度误差(MPJAE)从15.63°降至0.16°,平均关节位置误差(MPJPE)从18.51mm降至0.18mm,误差降低分别达98.98%和99.03%。实验表明,其他形态数据能提升目标形态的重建精度,验证了跨形态分词的优势;同时具备强零样本与少样本重建能力。
原文摘要 · Abstract (English)
Dexterous hands are essential for fine-grained manipulation, but their hardware designs vary substantially across embodiments. Differences in kinematics, joint definitions, and degrees of freedom make it difficult to define a shared state representation compared with parallel grippers. As a result, dexterous-hand data remains fragmented and difficult to use for joint training. In this work, we propose the Unified Dexterous Hand Model (UDHM), which maps human and robot hand states into a shared 22-DoF semantic interface. Based on UDHM, we introduce UniDexTok, a retargeting-free state tokenizer that learns embodiment-conditioned discrete tokens from standardized real joint states. UniDexTok provides a unified representation for heterogeneous dexterous hands without relying on retargeting or simulation data. Compared with the recent baseline UniHM, UniDexTok reduces MPJAE from 15.63 degrees to 0.16 degrees and MPJPE from 18.51 mm to 0.18 mm, corresponding to error reductions of 98.98% and 99.03%, respectively. These results improve reconstruction from centimeter-scale to sub-millimeter accuracy. Experiments further show that data from other embodiments improves target-embodiment reconstruction accuracy, demonstrating the benefit of cross-embodiment tokenization. UniDexTok also shows strong zero-shot and few-shot reconstruction ability when new dexterous hands are introduced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。