用人类操作视频训练机器人,大幅提升双手协作能力。
H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- 用人类第一视角操作视频预训练,学习自然抓取策略。
- 在仿真和真实场景中分别提升13.9%和40.5%性能。
- 适合需要高效学习双手操作的机器人研发人员。
机器人操作中的模仿学习面临一个根本挑战:高质量、大规模的机器人示范数据稀缺。现有机器人基础模型常通过跨形态机器人数据集预训练以扩大数据规模,但不同机器人形态与动作空间差异大,统一训练困难。本文提出H-RDT(Human to Robotics Diffusion Transformer),利用带有3D手部姿态标注的大规模第一视角人类操作视频,为机器人操作策略学习提供丰富的行为先验。我们采用两阶段训练范式:(1)在大规模人类操作数据上预训练;(2)在特定机器人数据上通过模块化动作编码器/解码器进行跨形态微调。基于含20亿参数的扩散变压器架构,采用流匹配建模复杂动作分布。大量仿真与真实世界实验表明,包括单任务、多任务、少样本学习及鲁棒性评估在内,H-RDT显著优于从零训练及现有先进方法(如Pi0和RDT),在仿真和真实环境中分别超越从零训练13.9%和40.5%。结果验证了人类操作数据可作为学习双手机器人操作策略的强大基础。
原文摘要 · Abstract (English)
Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data. Recent robotic foundation models often pre-train on cross-embodiment robot datasets to increase data scale, while they face significant limitations as the diverse morphologies and action spaces across different robot embodiments make unified training challenging. In this paper, we present H-RDT (Human to Robotics Diffusion Transformer), a novel approach that leverages human manipulation data to enhance robot manipulation capabilities. Our key insight is that large-scale egocentric human manipulation videos with paired 3D hand pose annotations provide rich behavioral priors that capture natural manipulation strategies and can benefit robotic policy learning. We introduce a two-stage training paradigm: (1) pre-training on large-scale egocentric human manipulation data, and (2) cross-embodiment fine-tuning on robot-specific data with modular action encoders and decoders. Built on a diffusion transformer architecture with 2B parameters, H-RDT uses flow matching to model complex action distributions. Extensive evaluations encompassing both simulation and real-world experiments, single-task and multitask scenarios, as well as few-shot learning and robustness assessments, demonstrate that H-RDT outperforms training from scratch and existing state-of-the-art methods, including Pi0 and RDT, achieving significant improvements of 13.9% and 40.5% over training from scratch in simulation and real-world experiments, respectively. The results validate our core hypothesis that human manipulation data can serve as a powerful foundation for learning bimanual robotic manipulation policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。