Kaiwu数据集整合多模态实时数据,助力机器人精细操作与人机协作研究。
Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction
- 构建包含20人、30类物体的多模态同步采集框架
- 收录11,664个操作实例,涵盖动作、压力、声音等10+维度信号
- 支持高精度标注与人机意图分析,适合机器人学习与交互研究
当前先进的机器人学习技术,如基础模型和人类示范学习,对大规模高质量数据集有巨大需求,而真实场景下的同步多模态数据仍是智能机器人领域的瓶颈。本文提出Kaiwu多模态数据集,解决复杂装配场景中动态信息与细粒度标注缺失的问题。该数据集集成20名受试者、30类交互物体,共收集11,664个完整操作实例,同步记录手部运动、操作压力、装配过程声音、多视角视频、高精度动作捕捉数据、第一视角眼动追踪视频及肌电图信号。每段演示均进行基于绝对时间戳的细粒度多级标注与语义分割标注。Kaiwu数据集旨在推动机器人学习、灵巧操作、人类意图解析及人机协作研究。
原文摘要 · Abstract (English)
Cutting-edge robot learning techniques including foundation models and imitation learning from humans all pose huge demands on large-scale and high-quality datasets which constitute one of the bottleneck in the general intelligent robot fields. This paper presents the Kaiwu multimodal dataset to address the missing real-world synchronized multimodal data problems in the sophisticated assembling scenario,especially with dynamics information and its fine-grained labelling. The dataset first provides an integration of human,environment and robot data collection framework with 20 subjects and 30 interaction objects resulting in totally 11,664 instances of integrated actions. For each of the demonstration,hand motions,operation pressures,sounds of the assembling process,multi-view videos, high-precision motion capture information,eye gaze with first-person videos,electromyography signals are all recorded. Fine-grained multi-level annotation based on absolute timestamp,and semantic segmentation labelling are performed. Kaiwu dataset aims to facilitate robot learning,dexterous manipulation,human intention investigation and human-robot collaboration research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。