arXiv:2409.15560cs.CVcs.HC2024-09被引 6

构建双人协作装配视觉数据集,助力机器人理解人类意图。

QUB-PHEO: A Visual-Based Dyadic Multi-View Dataset for Intention Inference in Collaborative Assembly

  • 采集70人完成36种子任务的双人协作视频,含面部、视线、手势等多模态标注。
  • 提供50人完整视频与70人视觉线索数据,支持机器学习模型训练。
  • 适合研究人机交互、意图识别与协作行为分析的学者使用。

QUB-PHEO 是一个基于视觉的双人多视角数据集,旨在推动装配作业中人机交互(HRI)与意图推断的研究。该数据集记录了70名参与者在36种不同子任务中的丰富多模态互动,其中一人作为‘机器人代理’。数据包含面部关键点、视线、手部动作、物体定位等详细视觉标注。提供两版数据:50人的完整视频与全部70人的视觉线索数据。该数据集有助于提升机器学习模型对细微交互线索和意图的理解能力,为相关领域研究提供支持。数据集将通过GitHub发布,需遵守最终用户许可协议(EULA)。

原文摘要 · Abstract (English)

QUB-PHEO introduces a visual-based, dyadic dataset with the potential of advancing human-robot interaction (HRI) research in assembly operations and intention inference. This dataset captures rich multimodal interactions between two participants, one acting as a 'robot surrogate,' across a variety of assembly tasks that are further broken down into 36 distinct subtasks. With rich visual annotations, such as facial landmarks, gaze, hand movements, object localization, and more for 70 participants, QUB-PHEO offers two versions: full video data for 50 participants and visual cues for all 70. Designed to improve machine learning models for HRI, QUB-PHEO enables deeper analysis of subtle interaction cues and intentions, promising contributions to the field. The dataset will be available at https://github.com/exponentialR/QUB-PHEO subject to an End-User License Agreement (EULA).

人机交互意图推断多模态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。