arXiv:2511.19684cs.CVcs.AI2025-11NeurIPS被引 4

构建工业场景下第一个多模态协作工作数据集,支持智能助手研究。

IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants

  • 采集3460段第一视角视频与1092段第三人称视频,覆盖多种工业任务。
  • 包含动作、错误标注及眼动、语音等多模态信息,支持复杂任务理解。
  • 适用于工业智能助手、人机协作与错误检测研究,挑战现有模型性能。

我们提出IndEgo,一个面向工业场景的多模态第一视角与第三人称数据集,涵盖装配/拆卸、物流与整理、检查与维修、木工等常见工业任务。数据集包含3,460段第一视角记录(约197小时),以及1,092段第三人称记录(约97小时)。重点聚焦双人协作任务,涉及高认知与体力负荷的操作。第一视角数据融合眼动、语音、声音、运动等丰富多模态信号,并提供详细标注(动作、摘要、错误标注、叙述)、元数据、处理输出(眼动、手部姿态、半稠密点云)及基准测试,涵盖流程性与非流程性任务理解、错误检测、基于推理的问题回答。基线评估表明,该数据集对当前先进多模态模型构成显著挑战。数据集已开源:https://huggingface.co/datasets/FraunhoferIPK/IndEgo。

原文摘要 · Abstract (English)

We introduce IndEgo, a multimodal egocentric and exocentric dataset addressing common industrial tasks, including assembly/disassembly, logistics and organisation, inspection and repair, woodworking, and others. The dataset contains 3,460 egocentric recordings (approximately 197 hours), along with 1,092 exocentric recordings (approximately 97 hours). A key focus of the dataset is collaborative work, where two workers jointly perform cognitively and physically intensive tasks. The egocentric recordings include rich multimodal data and added context via eye gaze, narration, sound, motion, and others. We provide detailed annotations (actions, summaries, mistake annotations, narrations), metadata, processed outputs (eye gaze, hand pose, semi-dense point cloud), and benchmarks on procedural and non-procedural task understanding, Mistake Detection, and reasoning-based Question Answering. Baseline evaluations for Mistake Detection, Question Answering and collaborative task understanding show that the dataset presents a challenge for the state-of-the-art multimodal models. Our dataset is available at: https://huggingface.co/datasets/FraunhoferIPK/IndEgo

工业智能多模态数据人机协作第一视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。