RDT2让机器人模型零样本适配新硬件,实现跨平台通用操作
RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization
- 用增强版通用操作接口收集超1万小时演示数据,构建大规模开放数据集
- 通过三阶段训练使模型在未见物体、场景、指令和机器人平台上实现零样本泛化
- 在打乒乓球等复杂任务中表现超越现有顶尖模型,适合具身智能研究者
视觉-语言-动作(VLA)模型虽有望实现通用机器人,但受限于数据稀缺、架构低效及跨硬件平台泛化能力不足。我们提出RDT2,一个基于70亿参数视觉语言模型的机器人基础模型,旨在实现开放词汇任务下对新型机器人平台的零样本部署。为此,我们利用改进的、与具体机器人无关的通用操作接口(UMI),收集了目前最大的开源机器人数据集之一——超过1万小时、涵盖多种机器人家族的示范数据。该方法采用新颖的三阶段训练流程,通过残差向量量化(RVQ)、流匹配和知识蒸馏,将离散语言知识与连续控制对齐,支持实时推理。结果表明,RDT2成为首个同时实现对未见物体、场景、指令及机器人平台零样本泛化的模型,并在灵巧操作、长时序任务和动态任务(如打乒乓球)中优于现有最优基线。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models hold promise for generalist robotics but currently struggle with data scarcity, architectural inefficiencies, and the inability to generalize across different hardware platforms. We introduce RDT2, a robotic foundation model built upon a 7B parameter VLM designed to enable zero-shot deployment on novel embodiments for open-vocabulary tasks. To achieve this, we collected one of the largest open-source robotic datasets--over 10,000 hours of demonstrations in diverse families--using an enhanced, embodiment-agnostic Universal Manipulation Interface (UMI). Our approach employs a novel three-stage training recipe that aligns discrete linguistic knowledge with continuous control via Residual Vector Quantization (RVQ), flow-matching, and distillation for real-time inference. Consequently, RDT2 becomes one of the first models that simultaneously zero-shot generalizes to unseen objects, scenes, instructions, and even robotic platforms. Besides, it outperforms state-of-the-art baselines in dexterous, long-horizon, and dynamic downstream tasks like playing table tennis. See https://rdt-robotics.github.io/rdt2/ for more information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。