用人类动作数据直接教会机器人新动作,提升抓取成功率。
MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation Policies
- 通过人机协同训练,直接迁移13项人类动作到机器人
- 9项任务零样本部署即达非平凡成功率,预训练微调提升40%
- 适合想用真实人类数据提升机器人操作能力的研究者
规模化真实机器人数据是模仿学习中的关键瓶颈,因此常依赖辅助数据进行策略训练。尽管图像或语言理解可借助网络数据集学习,但获取运动知识仍具挑战。人类数据因其丰富的操作行为多样性,成为宝贵资源。以往研究虽表明使用人类数据能提升鲁棒性和训练效率,但其最大潜力——使机器人策略直接学习新动作以完成任务——尚不明确。本文通过多任务人机协同训练系统性探索该可能。提出MotionTrans框架,包含数据采集系统、人类数据转换流程及加权协同训练策略。在30个任务上同时协同训练,成功将13个任务的人类动作直接迁移到可部署的端到端机器人策略中。值得注意的是,9项任务实现零样本成功。该方法还显著提升预训练-微调性能(+40%成功率)。消融实验揭示成功关键:与机器人数据协同训练及广泛覆盖相关动作。这些发现解锁了从人类数据中实现运动级学习的潜力,为机器人操作策略训练提供新思路。所有数据、代码和模型权重均已开源:https://motiontrans.github.io/
原文摘要 · Abstract (English)
Scaling real robot data is a key bottleneck in imitation learning, leading to the use of auxiliary data for policy training. While other aspects of robotic manipulation such as image or language understanding may be learned from internet-based datasets, acquiring motion knowledge remains challenging. Human data, with its rich diversity of manipulation behaviors, offers a valuable resource for this purpose. While previous works show that using human data can bring benefits, such as improving robustness and training efficiency, it remains unclear whether it can realize its greatest advantage: enabling robot policies to directly learn new motions for task completion. In this paper, we systematically explore this potential through multi-task human-robot cotraining. We introduce MotionTrans, a framework that includes a data collection system, a human data transformation pipeline, and a weighted cotraining strategy. By cotraining 30 human-robot tasks simultaneously, we direcly transfer motions of 13 tasks from human data to deployable end-to-end robot policies. Notably, 9 tasks achieve non-trivial success rates in zero-shot manner. MotionTrans also significantly enhances pretraining-finetuning performance (+40% success rate). Through ablation study, we also identify key factors for successful motion learning: cotraining with robot data and broad task-related motion coverage. These findings unlock the potential of motion-level learning from human data, offering insights into its effective use for training robotic manipulation policies. All data, code, and model weights are open-sourced https://motiontrans.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。