让机器人像人一样看、走、伸手,直接从真人动作数据学。
Hand-Eye Autonomous Delivery: Learning Humanoid Navigation, Locomotion and Reaching
- 分层设计:高层规划手眼目标,底层控制器跟踪真人动作数据。
- 在仿真和真实场景中均实现复杂环境下的导航与抓取。
- 解耦视觉感知与物理动作,便于扩展到新场景。
我们提出手眼自主配送(HEAD)框架,直接从人类动作和视觉感知数据中学习人形机器人的导航、运动及抓取技能。采用模块化设计,高层规划器指定人形机器人手部与眼部的目标位置和朝向,由底层策略控制全身运动。具体而言,底层全身体控制器基于大规模真人动捕数据,学习跟踪眼睛、左/右双手三个点的运动;高层策略则利用Aria眼镜采集的人类数据进行学习。该模块化方法将自我中心视觉感知与物理动作解耦,促进高效学习并具备向新场景扩展的能力。我们在仿真与真实世界中评估了该方法,验证了人形机器人在专为人类设计的复杂环境中完成导航与抓取的能力。
原文摘要 · Abstract (English)
We propose Hand-Eye Autonomous Delivery (HEAD), a framework that learns navigation, locomotion, and reaching skills for humanoids, directly from human motion and vision perception data. We take a modular approach where the high-level planner commands the target position and orientation of the hands and eyes of the humanoid, delivered by the low-level policy that controls the whole-body movements. Specifically, the low-level whole-body controller learns to track the three points (eyes, left hand, and right hand) from existing large-scale human motion capture data while high-level policy learns from human data collected by Aria glasses. Our modular approach decouples the ego-centric vision perception from physical actions, promoting efficient learning and scalability to novel scenes. We evaluate our method both in simulation and in the real-world, demonstrating humanoid's capabilities to navigate and reach in complex environments designed for humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。