用人类视频训练人形机器人,实现高质量动作生成与文本对齐。
Unveiling the Impact of Data and Model Scaling on High-Level Control for Humanoid Robots
- 构建自动化数据管道,生成260小时带语义标注的人形机器人运动数据
- 在数据与模型扩增下,动作重建误差降低37%,文本-动作对齐提升25%
- 可直接部署于真实人形机器人,适用于大规模动作学习研究
数据规模一直是机器人学习的关键瓶颈。对于人形机器人而言,人类视频与动作数据丰富且易获取,提供了免费的大规模数据来源,其动作语义有助于模态对齐与高层控制学习。然而,如何有效挖掘原始视频、提取可被机器人学习的表征,并用于可扩展的学习仍是一个开放问题。为此,我们提出Humanoid-Union,一个通过自动化流程生成的大规模数据集,包含超过260小时多样化的高质量人形机器人运动数据,语义标注源自人类动作视频,可通过相同流程持续扩展。基于此数据资源,我们设计了SCHUR——一个可扩展的学习框架,用于探索大规模数据对人形机器人高层控制的影响。实验表明,SCHUR在数据与模型扩增下实现了高精度的动作生成和强文本-动作对齐,相较于先前方法,MPJPE指标下重建性能提升37%,FID指标下对齐性能提升25%。其有效性也在真实人形机器人部署中得到验证。
原文摘要 · Abstract (English)
Data scaling has long remained a critical bottleneck in robot learning. For humanoid robots, human videos and motion data are abundant and widely available, offering a free and large-scale data source. Besides, the semantics related to the motions enable modality alignment and high-level robot control learning. However, how to effectively mine raw video, extract robot-learnable representations, and leverage them for scalable learning remains an open problem. To address this, we introduce Humanoid-Union, a large-scale dataset generated through an autonomous pipeline, comprising over 260 hours of diverse, high-quality humanoid robot motion data with semantic annotations derived from human motion videos. The dataset can be further expanded via the same pipeline. Building on this data resource, we propose SCHUR, a scalable learning framework designed to explore the impact of large-scale data on high-level control in humanoid robots. Experimental results demonstrate that SCHUR achieves high robot motion generation quality and strong text-motion alignment under data and model scaling, with 37\% reconstruction improvement under MPJPE and 25\% alignment improvement under FID comparing with previous methods. Its effectiveness is further validated through deployment in real-world humanoid robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。