构建首个统一平台的多机器人形态操作数据集,支持高效学习与泛化。
RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
- 基于人类遥控采集107万条轨迹,覆盖96类物体和4种机器人形态。
- 包含5000条真实失败案例及原因标注,支持策略纠错与鲁棒训练。
- 适配视觉语言动作模型,助力多任务场景下的高成功率泛化。
本文提出RoboMIND(机器人操作多形态智能基准数据集),包含跨479个多样化任务、96类物体的107,000条示范轨迹。数据通过人类遥控采集,涵盖多视角观测、本体感知状态及语言任务描述,统一在标准化平台完成,覆盖Franka Emika Panda、UR5e、AgileX双臂机器人及双灵巧手人形机器人四种形态。数据集还包含5,000条真实世界失败示范,每条附详细原因,支持策略反思与修正。我们构建了Isaac Sim中的数字孪生环境,可低成本生成额外训练数据并高效评估。实验表明,利用RoboMIND,VLA模型在单任务与多任务设置下均实现高成功率与强泛化能力。据我们所知,RoboMIND是目前规模最大的统一平台多形态遥操作数据集,为机器人训练提供大规模高质量数据支持。
原文摘要 · Abstract (English)
In this paper, we introduce RoboMIND (Multi-embodiment Intelligence Normative Data for Robot Manipulation), a dataset containing 107k demonstration trajectories across 479 diverse tasks involving 96 object classes. RoboMIND is collected through human teleoperation and encompasses comprehensive robotic-related information, including multi-view observations, proprioceptive robot state information, and linguistic task descriptions. To ensure data consistency and reliability for imitation learning, RoboMIND is built on a unified data collection platform and a standardized protocol, covering four distinct robotic embodiments: the Franka Emika Panda, the UR5e, the AgileX dual-arm robot, and a humanoid robot with dual dexterous hands. Our dataset also includes 5k real-world failure demonstrations, each accompanied by detailed causes, enabling failure reflection and correction during policy learning. Additionally, we created a digital twin environment in the Isaac Sim simulator, replicating the real-world tasks and assets, which facilitates the low-cost collection of additional training data and enables efficient evaluation. To demonstrate the quality and diversity of our dataset, we conducted extensive experiments using various imitation learning methods for single-task settings and state-of-the-art Vision-Language-Action (VLA) models for multi-task scenarios. By leveraging RoboMIND, the VLA models achieved high manipulation success rates and demonstrated strong generalization capabilities. To the best of our knowledge, RoboMIND is the largest multi-embodiment teleoperation dataset collected on a unified platform, providing large-scale and high-quality robotic training data. Our project is at https://x-humanoid-robomind.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。