从人类视频中学习通用交互技能,让机器人零样本复现复杂动作
HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos
- 用视频生成物理合理的机器人交互数据,支持大规模增强
- 仅凭单个视频示范就学会10种复杂技能,零样本迁移到真实机器人
- 适合想快速构建通用人形机器人交互能力的研究者
让类人机器人完成敏捷、自适应的交互任务是机器人领域长期挑战。现有方法受限于真实交互数据稀缺或需繁琐的任务特定奖励设计,难以扩展。为此,我们提出HumanX——一个全栈式框架,将人类视频转化为可泛化的类人交互技能,无需任务特定奖励。HumanX包含两个协同设计组件:XGen,一种从视频中合成多样化且物理合理的机器人交互数据的数据生成管道,支持可扩展的数据增强;XMimic,一种统一的模仿学习框架,用于学习通用交互技能。在篮球、足球、羽毛球、货物搬运和反应式格斗五个不同领域上评估,HumanX成功习得10种不同技能,并零样本迁移至真实Unitree G1类人机器人。所学能力包括无需外部感知的假动作转身跳投等复杂动作,以及连续10轮以上的人机传球序列——全部基于单个视频示范。实验表明,HumanX的泛化成功率超过先前方法8倍以上,展示了可扩展、任务无关的学习路径,用于获取多样化的现实世界机器人交互技能。
原文摘要 · Abstract (English)
Enabling humanoid robots to perform agile and adaptive interactive tasks has long been a core challenge in robotics. Current approaches are bottlenecked by either the scarcity of realistic interaction data or the need for meticulous, task-specific reward engineering, which limits their scalability. To narrow this gap, we present HumanX, a full-stack framework that compiles human video into generalizable, real-world interaction skills for humanoids, without task-specific rewards. HumanX integrates two co-designed components: XGen, a data generation pipeline that synthesizes diverse and physically plausible robot interaction data from video while supporting scalable data augmentation; and XMimic, a unified imitation learning framework that learns generalizable interaction skills. Evaluated across five distinct domains--basketball, football, badminton, cargo pickup, and reactive fighting--HumanX successfully acquires 10 different skills and transfers them zero-shot to a physical Unitree G1 humanoid. The learned capabilities include complex maneuvers such as pump-fake turnaround fadeaway jumpshots without any external perception, as well as interactive tasks like sustained human-robot passing sequences over 10 consecutive cycles--learned from a single video demonstration. Our experiments show that HumanX achieves over 8 times higher generalization success than prior methods, demonstrating a scalable and task-agnostic pathway for learning versatile, real-world robot interactive skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。