用人类视频训练手部精细操作模型,提升机器人动作预测能力
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
- 通过人体视角视频提取手指关键点作为动作输入
- 在900小时数据上训练,零样本迁移性能超扩散策略50%以上
- 引入手部一致性损失,精准建模手指姿态与物体交互
建模灵巧手-物体交互极具挑战,需理解细微指部运动如何通过接触影响环境。现有世界模型多依赖粗粒度动作空间,难以捕捉精细操作。为此,我们提出DexWM,一种基于过去状态和灵巧动作预测未来环境隐状态的世界模型。为应对精细标注数据稀缺问题,DexWM利用人体视角视频中提取的手指关键点表示动作,实现对超过900小时的人类及非灵巧机器人数据的训练。此外,为准确建模灵巧性,我们发现仅预测视觉特征不足;因此引入辅助手部一致性损失,强制保持手部构型准确性。DexWM在未来状态预测上优于基于文本、导航或全身动作的先前世界模型,并在Franka Panda机械臂搭配Allegro夹具上展现出强大零样本迁移能力,在抓取、放置和伸展任务中平均超越Diffusion Policy超过50%。
原文摘要 · Abstract (English)
Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction modeling, they typically rely on coarse action spaces that fail to capture fine-grained dexterity. We, therefore, introduce DexWM, a Dexterous Interaction World Model that predicts future latent states of the environment conditioned on past states and dexterous actions. To overcome the scarcity of finely annotated dexterous datasets, DexWM represents actions using finger keypoints extracted from egocentric videos, enabling training on over 900 hours of human and non-dexterous robot data. Further, to accurately model dexterity, we find that predicting visual features alone is insufficient; therefore, we incorporate an auxiliary hand consistency loss that enforces accurate hand configurations. DexWM outperforms prior world models conditioned on text, navigation, or full-body actions in future-state prediction and demonstrates strong zero-shot transfer to unseen skills on a Franka Panda arm with an Allegro gripper, surpassing Diffusion Policy by over 50% on average across grasping, placing, and reaching tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。