arXiv:2604.08342cs.LG2026-04

用眼球追踪数据生成问题,让AR视频理解更贴近真实人类行为。

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment

论文配图:EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment
图 1 · 摘自论文原文
  • 基于眼球注视数据生成问题,反映真实注意力分布。
  • 涵盖超5000个问答对,覆盖100多小时视频内容。
  • 适合研究长时序视频理解与AR人机交互的学者。

长时序第一视角视频理解近年来受到广泛关注,增强现实(AR)是其重要应用领域之一。然而,该任务仍具挑战性,需对长时间上下文进行推理,并处理多样且无结构的活动。尽管已有多个基准数据集,但多数第一视角数据集依赖佩戴式相机,主要关注视觉内容,较少考虑用户在生成视频相关问题时的行为。EgoEverything通过利用从眼动数据中抽象出的人类注意力信号来生成问题,显式地纳入人类行为因素。该基准包含超过5000个多项选择题-答案对,覆盖100多小时的视频内容。通过在问题生成中整合人类注意力信号,它更真实地反映了自然人类行为,为AR环境下的长时序第一视角视频理解提供了更真实的评估场景。

原文摘要 · Abstract (English)

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging due to the need for reasoning over extended temporal contexts and diverse, unstructured activities. Although several benchmarks exist, most egocentric datasets rely on human worn cameras and focus mainly on visual content, with limited consideration of underlying user behavior when forming video-related queries. EgoEverything is a benchmark that explicitly considers human behavior by leveraging human attention signals, abstracted from gaze data, when generating questions. It comprises over 5,000 multiple choice question answer pairs, spanning more than 100 hours of video. By integrating human attention signals during question generation, it more faithfully captures natural human behavior and offers a realistic evaluation setting for long-context egocentric video understanding in AR.

视频理解AR眼动追踪长时序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。