用人类视频训练机器人,减少80%的实机演示需求
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
- 通过双阶段特征蒸馏对齐人与机器人的视觉表征
- 人类数据可替代80%的机器人实拍示范数据
- 适合希望降低机器人训练成本的研究者
扩大机器人学习能力受限于机器人示范数据稀缺,而人类视频提供了海量未被利用的交互数据。然而,人类手部与机器人臂之间的形态差异仍是关键挑战。现有跨形态迁移方法多依赖视觉编辑,常因视觉外观和三维几何的固有差异引入视觉伪影。为此,我们提出LIDEA(隐式特征蒸馏与显式几何对齐)框架,使策略学习受益于人类示范。在二维视觉域,采用双阶段传递蒸馏管道,在共享潜在空间中对齐人与机器人表征;在三维几何域,提出一种与形态无关的对齐策略,显式解耦形态与交互几何,确保一致的三维感知。大量实验验证了LIDEA在数据效率和分布外鲁棒性方面的有效性:人类数据可替代高达80%的高成本机器人示范,且框架成功将人类视频中的未见模式迁移至分布外场景,实现泛化。
原文摘要 · Abstract (English)
Scaling up robot learning is hindered by the scarcity of robotic demonstrations, whereas human videos offer a vast, untapped source of interaction data. However, bridging the embodiment gap between human hands and robot arms remains a critical challenge. Existing cross-embodiment transfer strategies typically rely on visual editing, but they often introduce visual artifacts due to intrinsic discrepancies in visual appearance and 3D geometry. To address these limitations, we introduce LIDEA (Implicit Feature Distillation and Explicit Geometric Alignment), an imitation learning framework in which policy learning benefits from human demonstrations. In the 2D visual domain, LIDEA employs a dual-stage transitive distillation pipeline that aligns human and robot representations in a shared latent space. In the 3D geometric domain, we propose an embodiment-agnostic alignment strategy that explicitly decouples embodiment from interaction geometry, ensuring consistent 3D-aware perception. Extensive experiments empirically validate LIDEA from two perspectives: data efficiency and OOD robustness. Results show that human data substitutes up to 80% of costly robot demonstrations, and the framework successfully transfers unseen patterns from human videos for out-of-distribution generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。