用人体与机器人肢体对应关系,对齐视觉表征,让机器人更好学习人类演示。
LACE: Latent Visual Representation for Cross-Embodiment Learning

- 利用人体与机器人共享肢体的对应关系,作为稀疏监督对齐视觉特征。
- 仅需单次机器人示范,零样本迁移性能比基线高65%。
- 适合数据稀缺、跨机器人形态的强化学习场景。
从人类示范中进行跨形态学习受人类与机器人形态间视觉差异的阻碍。尽管自监督学习(SSL)骨干网络能编码通用物体的丰富类别语义,但它们无法建立人手与机器人手之间的对应关系。我们提出LACE框架,通过利用不同形态间共享身体部位的对应关系作为稀疏监督,在这些骨干网络的潜在空间中对齐人类与机器人的视觉表示。这些标注可通过正向运动学自动获取,且单次机器人示范即可完成模型训练。我们的语义对齐损失匹配对应特征产生的分布,将像素级监督提升至语义级对齐;同时,格拉姆损失保留了预训练特征的质量。该对齐使机器人策略能在机器人示范稀缺时利用大量人类数据:在零样本迁移中,使用LACE-DINO的策略相比DINO显著提升65%,在低数据和分布外环境中也保持一致优势。
原文摘要 · Abstract (English)
Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich inter-class semantics of general objects, we show they fail to establish correspondence between human and robot hands. We propose LACE, a framework that aligns human and robot visual representations in the latent space of these backbones by leveraging correspondences between shared body parts across embodiments as sparse supervision. These annotations can be automatically obtained via forward kinematics, and single robot demonstration is sufficient to train the model. Our semantic alignment loss matches distributions incurred by corresponding features, lifting patch-level supervision to semantic-level alignment, while a Gram loss preserves pretrained feature quality. This alignment enables robot policies to leverage abundant human data when robot demonstrations are scarce: in zero-shot transfer, policies using LACE-DINO outperform those using DINO by a large margin (65\%), with consistent gains in low-data regimes and out-of-distribution environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。