通过人体示范学习灵巧操作,实现跨机器人零样本迁移。
AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning

- 构建统一动作空间,对齐人手与机器人的操作表示。
- 结合对抗学习,让视觉特征摆脱特定机器人外观影响。
- 仅需少量数据即可在真实场景中完成灵巧操作任务。
灵巧操作是具身智能的基础能力,但其规模化面临挑战:机器人示范数据采集成本高,且不同机器人形态导致动作空间差异。基于异构数据训练的策略会将任务相关视觉线索与特定形态外观混淆,限制跨形态泛化能力。本文提出AdvDex,一种统一的视觉-语言-动作框架,可从人体与机器人示范中学习灵巧操作。首先,构建OmniShare——大规模多模态人体操作示范数据集,提供高质量运动学监督与触觉测量,减少对机器人遥操作的依赖。其次,提出联合对齐动作空间(JAAS),以$ m{SE}(3)$手腕位姿和15个指关节构成标准动作表示,实现人体手、灵巧机械手与平行夹爪的功能对齐。最后,采用领域对抗学习降低视觉表征中的形态特异性信息。在手势预测与真实世界灵巧操作实验中,相比基线模型表现更优,实现了有效的零样本人体到机器人技能迁移,对未见物体与环境具备泛化能力,并支持数据高效的少样本适应。
原文摘要 · Abstract (English)
Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spaces vary across embodiments. Policies trained on heterogeneous data can also entangle task-relevant visual cues with embodiment-specific appearance, limiting cross-embodiment generalization. We present AdvDex, a unified Vision-Language-Action framework for learning dexterous manipulation from human and robot demonstrations. First, we introduce OmniShare, a large-scale multimodal dataset of human manipulation demonstrations that provides high-quality kinematic supervision and tactile measurements while reducing reliance on robot teleoperation. Second, we propose the Joint-Aligned Action Space (JAAS), a canonical action representation comprising an $\mathrm{SE}(3)$ wrist pose and 15 finger joints, thereby functionally aligning human hands, dexterous robot hands, and parallel grippers. Finally, we use domain-adversarial learning to reduce embodiment-specific information in the learned visual representation. Experiments on hand-action prediction and real-world dexterous manipulation show consistent improvements over baselines, effective zero-shot human-to-robot skill transfer, generalization to unseen objects and environments, and data-efficient few-shot adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。