构建首个融合力觉与分层标注的大规模移动操作数据集
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
- 采集真实机器人操作数据,同步记录视觉、关节、六维力矩信号
- 包含25,469个任务片段(约94小时),覆盖长时序接触式操作
- 提供子目标与基础动作双层标注,适合训练通用智能体
随着机器人从受控环境进入非结构化人类场景,实现能可靠执行自然语言指令的通用智能体仍是核心挑战。稳健的移动操作进展需要大规模多模态数据集,涵盖高接触频率和长时程任务,但现有资源普遍缺乏同步的力-扭矩传感、分层标注和明确失败案例。我们提出AIRoA MoMa数据集,一个大规模真实世界移动操作多模态数据集。包含同步的RGB图像、关节状态、六轴腕部力-扭矩信号及内部机器人状态,并引入新颖的两级标注体系(子目标与基础动作),支持层次化学习与错误分析。初始版本基于人机协作机器人(HSR)采集,共25,469个任务片段(约94小时),已标准化为LeRobot v2.1格式。通过整合移动操作、高接触交互与长时程结构,AIRoA MoMa为下一代视觉-语言-动作模型提供了关键基准。首个版本现已在https://huggingface.co/datasets/airoa-org/airoa-moma 发布。
原文摘要 · Abstract (English)
As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical annotations, and explicit failure cases. We address this gap with the AIRoA MoMa Dataset, a large-scale real-world multimodal dataset for mobile manipulation. It includes synchronized RGB images, joint states, six-axis wrist force-torque signals, and internal robot states, together with a novel two-layer annotation schema of sub-goals and primitive actions for hierarchical learning and error analysis. The initial dataset comprises 25,469 episodes (approx. 94 hours) collected with the Human Support Robot (HSR) and is fully standardized in the LeRobot v2.1 format. By uniquely integrating mobile manipulation, contact-rich interaction, and long-horizon structure, AIRoA MoMa provides a critical benchmark for advancing the next generation of Vision-Language-Action models. The first version of our dataset is now available at https://huggingface.co/datasets/airoa-org/airoa-moma .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。