构建首个基于流程文本的自拍视角错误动作数据集,助力智能纠错系统研发。
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
- 基于多领域流程文本录制自拍视频,标注动作与文本对齐关系
- 引入文本信息后,错误检测准确率显著提升,证明文本不可或缺
- 适合研究人机交互、智能辅导与工业安全的学者使用
错误动作检测对于开发能够识别工人失误并提供反馈的智能档案系统至关重要。现有研究多关注自由活动中的视觉明显错误,采用纯视频方法进行检测。但在遵循流程文本的任务中,模型仅凭视觉难以判断动作是否正确。此外,当前错误数据集极少使用流程文本进行视频采集,除烹饪外几乎空白。为此,本文提出 EgoOops 数据集,通过自拍视角视频记录在不同领域中遵循流程文本时发生的错误行为,包含三类标注:视频-文本对齐、错误标签及错误描述。我们还提出一种结合视频-文本对齐与错误分类的方法,以利用文本信息。实验表明,引入流程文本对错误检测至关重要。数据集可通过 https://y-haneji.github.io/EgoOops-project-page/ 获取。
原文摘要 · Abstract (English)
Mistake action detection is crucial for developing intelligent archives that detect workers' errors and provide feedback. Existing studies have focused on visually apparent mistakes in free-style activities, resulting in video-only approaches to mistake detection. However, in text-following activities, models cannot determine the correctness of some actions without referring to the texts. Additionally, current mistake datasets rarely use procedural texts for video recording except for cooking. To fill these gaps, this paper proposes the EgoOops dataset, where egocentric videos record erroneous activities when following procedural texts across diverse domains. It features three types of annotations: video-text alignment, mistake labels, and descriptions for mistakes. We also propose a mistake detection approach, combining video-text alignment and mistake label classification to leverage the texts. Our experimental results show that incorporating procedural texts is essential for mistake detection. Data is available through https://y-haneji.github.io/EgoOops-project-page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。