让机器人通过人类动作学习物理规律,实现跨形态高效操作。
LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation

- 在共享潜空间中对齐人手与机器人的动作动态,超越几何模仿。
- 仅用20分钟人手数据,实现超过90%的操作准确率。
- 低成本手套采集高精度动作数据,适合各类机械手使用。
人类手部示范为机器人学习提供了直接且可扩展的物理交互数据源。尽管手动重定向对于建立不同形态间的运动学对应关系至关重要,但稳健迁移还需超越几何层面,解决人与机器人操作间物理动力学的底层对齐问题。为此,我们提出LaST-HD,一种新型的人类到机器人的动作学习范式,通过在共享潜推理空间中对齐人类手部与机器人示范,扩展了先推理后执行的视觉-语言-动作模型(VLA)。LaST-HD不模仿人类运动学,而是训练一个辅助的动作条件世界模型,基于未配对的人类手部与机器人轨迹生成统一的潜目标。在共享前向动力学空间中对齐跨体态表示后,这些目标监督LaST-HD的潜推理过程,使其内化共享的物理动态,从而驱动高效的人类动作学习。此外,我们开发了低成本的“离实验室”(OOL)手套,专为LaST-HD设计,用于采集人类手部数据。所捕获数据提供精确关键点,并作为跨夹爪和灵巧手的通用动作监督信号。结合对齐的潜空间与高保真人类手部数据,我们构建了一种渐进式混合到人类的训练方案,包括混合人机协同训练与训练后在线修正。通过混合协同训练,LaST-HD仅依赖人类手部示范即可提升对新物体、新场景和新位置的泛化能力。经在线修正后,该模型进一步适应新环境,仅用20分钟的OOL手套数据即实现超90%的准确率。
原文摘要 · Abstract (English)
Human-hand demonstrations provide a direct and scalable source of physical interaction data for robot learning. While manual retargeting is indispensable for establishing kinematic action correspondence across different morphologies, robust transfer requires going beyond geometry to address the underlying alignment of physical dynamics between human and robot manipulation. To address this, we introduce LaST-HD, a novel human-to-robot action learning paradigm that extends reasoning-before-acting VLA by aligning human-hand and robot demonstrations in a shared latent reasoning space. Rather than mimicking human kinematics, LaST-HD trains an auxiliary action-conditioned world model on unpaired human-hand and robot trajectories to synthesize unified latent targets. After aligning cross-embodiment representations in this shared forward-dynamics space, these targets supervise LaST-HD's latent reasoning process, enabling it to internalize shared physical dynamics and drive efficient human-hand action learning. Moreover, we develop Out-of-Lab (OOL) Glove, a low-cost motion-capture glove tailored to LaST-HD for human-hand data collection. The captured human data provide precise keypoints and serve as universal action supervision across grippers and dexterous hands. Armed with the aligned latent space and high-fidelity human-hand data, we develop a progressive mixed-to-human training recipe comprising mixed human-robot co-training and human-hand online correction post-training. Through mixed co-training, LaST-HD improves generalization to novel objects, scenes, and positions using only human-hand demonstrations. With online correction, LaST-HD further adapts to novel environments and achieves over 90\% accuracy using only 20 minutes of OOL glove data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。