arXiv:2502.04144cs.CV2025-02CVPR被引 105

构建首个真实厨房场景的高精度第一视角视频数据集,支持多模态细粒度分析。

HD-EPIC: A Highly-Detailed Egocentric Video Dataset

论文配图:HD-EPIC: A Highly-Detailed Egocentric Video Dataset
图 1 · 摘自论文原文
  • 通过数字孪生与注视追踪实现3D空间标注,覆盖动作、食材、音频等多维度信息
  • 41小时视频含59000个细粒度动作、20000次物体移动,平均每分钟263个标注
  • 适用于视觉问答、动作识别与长时序分割,挑战现有视觉语言模型性能极限

我们提出一个新采集的基于厨房的第一视角视频验证数据集,手动标注了高度详细且相互关联的真实标签,涵盖食谱步骤、细粒度动作、带营养值的食材、移动物体及音频标注。所有标注均通过场景、固定装置、物体位置的数字孪生实现3D定位,并结合注视信息进行预训练。视频来自多样化家庭环境中的非脚本化录制,使HD-EPIC成为首个在自然环境中收集但具备实验室级精细标注的数据集。我们通过一项包含26000个问题的挑战性视觉问答基准,评估其识别食谱、食材、营养、细粒度动作、3D感知、物体运动和注视方向的能力。当前最强的长上下文模型Gemini Pro仅达到38.5%准确率,凸显该数据集难度与现有视觉语言模型的不足。此外,我们在HD-EPIC上评估了动作识别、声音识别与长期视频-物体分割任务。该数据集包含41小时视频,覆盖9个厨房,拥有413个厨房设施的数字孪生,记录69道菜谱,59000个细粒度动作,51000个音频事件,20000次物体移动以及37000个提升至3D的物体掩码。平均每分钟视频有263个标注。

原文摘要 · Abstract (English)

We present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, all annotations are grounded in 3D through digital twinning of the scene, fixtures, object locations, and primed with gaze. Footage is collected from unscripted recordings in diverse home environments, making HDEPIC the first dataset collected in-the-wild but with detailed annotations matching those in controlled lab environments. We show the potential of our highly-detailed annotations through a challenging VQA benchmark of 26K questions assessing the capability to recognise recipes, ingredients, nutrition, fine-grained actions, 3D perception, object motion, and gaze direction. The powerful long-context Gemini Pro only achieves 38.5% on this benchmark, showcasing its difficulty and highlighting shortcomings in current VLMs. We additionally assess action recognition, sound recognition, and long-term video-object segmentation on HD-EPIC. HD-EPIC is 41 hours of video in 9 kitchens with digital twins of 413 kitchen fixtures, capturing 69 recipes, 59K fine-grained actions, 51K audio events, 20K object movements and 37K object masks lifted to 3D. On average, we have 263 annotations per minute of our unscripted videos.

第一视角视频标注数字孪生视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。