用2000万条人类视频训练机器人,让其听懂指令就能动作
Learning from Massive Human Videos for Universal Humanoid Pose Control
- 从网络视频提取动作数据,转为机器人可学的文本-动作对
- 训练出的UH-1模型能根据文字指令控制机器人完成多样化动作
- 在仿真和真实世界中均表现良好,适合做通用人形机器人控制
可扩展的人形机器人学习对其在现实场景中的部署至关重要。传统方法主要依赖强化学习或遥操作实现全身控制,但受限于模拟环境多样性及示范数据收集成本高昂。相比之下,人类视频广泛存在,蕴含丰富的语义与运动信息,可显著提升人形机器人的泛化能力。本文提出Humanoid-X,一个包含超过2000万个人形机器人姿态及其对应文本动作描述的大规模数据集,旨在利用这一丰富数据源。该数据集通过综合流程构建:互联网数据挖掘、视频字幕生成、人类动作向人形机器人迁移以及面向真实部署的策略学习。基于Humanoid-X,我们进一步训练了大型人形模型UH-1,该模型以文本指令为输入,输出相应动作以控制人形机器人。大量仿真与真实世界实验验证了该可扩展训练方法在基于文本的人形机器人控制中具备优异泛化能力,标志着迈向适应性强、具备现实部署能力的人形机器人的重要一步。
原文摘要 · Abstract (English)
Scalable learning of humanoid robots is crucial for their deployment in real-world applications. While traditional approaches primarily rely on reinforcement learning or teleoperation to achieve whole-body control, they are often limited by the diversity of simulated environments and the high costs of demonstration collection. In contrast, human videos are ubiquitous and present an untapped source of semantic and motion information that could significantly enhance the generalization capabilities of humanoid robots. This paper introduces Humanoid-X, a large-scale dataset of over 20 million humanoid robot poses with corresponding text-based motion descriptions, designed to leverage this abundant data. Humanoid-X is curated through a comprehensive pipeline: data mining from the Internet, video caption generation, motion retargeting of humans to humanoid robots, and policy learning for real-world deployment. With Humanoid-X, we further train a large humanoid model, UH-1, which takes text instructions as input and outputs corresponding actions to control a humanoid robot. Extensive simulated and real-world experiments validate that our scalable training approach leads to superior generalization in text-based humanoid control, marking a significant step toward adaptable, real-world-ready humanoid robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。