首个12亿参数的双臂机器人扩散模型,实现零样本泛化与少样本学习。
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

- 采用扩散模型与可扩展Transformer,处理双臂动作的多模态与高频率特性。
- 在6000+双臂任务数据上微调,支持语言指令理解与1~5次演示学新技能。
- 统一物理可解释动作空间,提升跨机器人迁移能力,适合具身智能研究者。
双臂操作对机器人至关重要,但构建基础模型极具挑战,原因在于双臂协调的固有复杂性(导致多模态动作分布)以及训练数据稀缺。本文提出机器人扩散变换器(RDT),首个用于双臂操作的扩散基础模型。RDT利用扩散模型有效表示多模态动作,创新设计可扩展的Transformer以处理多模态输入异质性,并捕捉机器人数据的非线性与高频特征。为缓解数据稀缺问题,我们引入物理可解释统一动作空间,统一不同机器人的动作表示,同时保留原始动作的物理意义,促进可迁移物理知识的学习。基于此设计,我们在迄今最大的多机器人数据集上预训练了12亿参数的RDT,是当前最大的基于扩散模型的机器人操作基础模型。最后,我们在自建的包含6000+任务片段的多任务双臂数据集上进行微调,以强化其操作能力。实机实验表明,RDT显著优于现有方法:具备未见物体与场景的零样本泛化能力,能理解并遵循语言指令,仅需1~5个示范即可学会新技能,并有效完成复杂灵巧任务。
原文摘要 · Abstract (English)
Bimanual manipulation is essential in robotics, yet developing foundation models is extremely challenging due to the inherent complexity of coordinating two robot arms (leading to multi-modal action distributions) and the scarcity of training data. In this paper, we present the Robotics Diffusion Transformer (RDT), a pioneering diffusion foundation model for bimanual manipulation. RDT builds on diffusion models to effectively represent multi-modality, with innovative designs of a scalable Transformer to deal with the heterogeneity of multi-modal inputs and to capture the nonlinearity and high frequency of robotic data. To address data scarcity, we further introduce a Physically Interpretable Unified Action Space, which can unify the action representations of various robots while preserving the physical meanings of original actions, facilitating learning transferrable physical knowledge. With these designs, we managed to pre-train RDT on the largest collection of multi-robot datasets to date and scaled it up to 1.2B parameters, which is the largest diffusion-based foundation model for robotic manipulation. We finally fine-tuned RDT on a self-created multi-task bimanual dataset with over 6K+ episodes to refine its manipulation capabilities. Experiments on real robots demonstrate that RDT significantly outperforms existing methods. It exhibits zero-shot generalization to unseen objects and scenes, understands and follows language instructions, learns new skills with just 1~5 demonstrations, and effectively handles complex, dexterous tasks. We refer to https://rdt-robotics.github.io/rdt-robotics/ for the code and videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。