用生成式数字孪生构建双臂机器人训练与评估平台,提升操作成功率。
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

- 基于2D图像生成可交互的3D数字孪生,模拟多样真实场景。
- 通过语言模型生成带空间约束的精准机械臂控制代码,任务成功率提升超70%。
- 适合研究双臂协作、仿真训练与真实部署对齐的机器人开发者。
在快速发展的机器人领域,双臂协同与复杂物体操作是实现高级自主系统的关键能力。然而,高质量演示数据稀缺和真实世界对齐的评估基准缺乏严重制约了该方向的发展。为此,我们提出RoboTwin——一种基于3D生成基础模型和大语言模型的生成式数字孪生框架,能够生成多样化专家数据集,并提供与真实世界对齐的双臂机器人任务评估平台。RoboTwin可从单张2D图像生成物体的多样化数字孪生,构建逼真且可交互的场景;同时引入空间关系感知的代码生成框架,结合物体标注与大语言模型,分解任务、确定空间约束并生成精确的机器人运动代码。该框架提供包含仿真与真实数据的综合性评估基准,实现训练与实际表现的更好对齐。我们在开源COBOT Magic Robot平台上验证了该方法:在少量真实样本微调下,基于RoboTwin生成数据预训练的策略,相比仅依赖真实数据训练的模型,在单臂任务上成功率提升超过70%,双臂任务提升超过40%。
原文摘要 · Abstract (English)
In the rapidly advancing field of robotics, dual-arm coordination and complex object manipulation are essential capabilities for developing advanced autonomous systems. However, the scarcity of diverse, high-quality demonstration data and real-world-aligned evaluation benchmarks severely limits such development. To address this, we introduce RoboTwin, a generative digital twin framework that uses 3D generative foundation models and large language models to produce diverse expert datasets and provide a real-world-aligned evaluation platform for dual-arm robotic tasks. Specifically, RoboTwin creates varied digital twins of objects from single 2D images, generating realistic and interactive scenarios. It also introduces a spatial relation-aware code generation framework that combines object annotations with large language models to break down tasks, determine spatial constraints, and generate precise robotic movement code. Our framework offers a comprehensive benchmark with both simulated and real-world data, enabling standardized evaluation and better alignment between simulated training and real-world performance. We validated our approach using the open-source COBOT Magic Robot platform. Policies pre-trained on RoboTwin-generated data and fine-tuned with limited real-world samples demonstrate significant potential for enhancing dual-arm robotic manipulation systems by improving success rates by over 70% for single-arm tasks and over 40% for dual-arm tasks compared to models trained solely on real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。