从真人视频学动作,让机器人像人一样操作物体
HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos
- 从单目视频提取人体与物体轨迹,构建结构化动作数据集
- 训练出可零样本部署的强化学习策略,在真实机器人上完成6项任务
- 无需额外标注,适合快速获取复杂人机交互技能
由于运动数据稀缺且接触频繁,实现稳健的全身人形机器人-物体交互仍具挑战。我们提出HDMI(HumanoiD iMitation for Interaction),一种简单通用的框架,直接从单目RGB视频中学习全身人形机器人-物体交互技能。该流程包括:(i) 从非受限视频中提取并重定向人体与物体轨迹,构建结构化运动数据集;(ii) 训练强化学习(RL)策略以协同追踪机器人与物体状态,采用三项关键设计:统一物体表示、残差动作空间和通用交互奖励;(iii) 零样本部署于真实人形机器人。在Unitree G1人形机器人上开展大量仿真到现实实验,验证了方法的鲁棒性与泛化性:HDMI实现67次连续过门,真实世界成功完成6项不同运动操作任务,仿真中完成14项任务。结果表明,HDMI是一种从真人视频获取交互式人形技能的简洁通用方案。
原文摘要 · Abstract (English)
Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for Interaction), a simple and general framework that learns whole-body humanoid-object interaction skills directly from monocular RGB videos. Our pipeline (i) extracts and retargets human and object trajectories from unconstrained videos to build structured motion datasets, (ii) trains a reinforcement learning (RL) policy to co-track robot and object states with three key designs: a unified object representation, a residual action space, and a general interaction reward, and (iii) zero-shot deploys the RL policies on real humanoid robots. Extensive sim-to-real experiments on a Unitree G1 humanoid demonstrate the robustness and generality of our approach: HDMI achieves 67 consecutive door traversals and successfully performs 6 distinct loco-manipulation tasks in the real world and 14 tasks in simulation. Our results establish HDMI as a simple and general framework for acquiring interactive humanoid skills from human videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。