arXiv:2607.02096cs.CV2026-07

构建首个长时第一人称视频指代理解基准,挑战模型在45分钟长视频中定位稀疏目标。

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

论文配图:LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
图 1 · 摘自论文原文
  • 基于Ego4D数据集构建长时第一人称视频指代基准,平均时长45分钟
  • 包含1498个指代表达,目标出现极稀疏且场景复杂多变
  • 现有模型在该任务上表现显著下降,凸显长时视频理解瓶颈

第一人称视频记录了丰富多样的人-物交互行为,是理解与物体相关人类活动的基础资源。视频指代表达理解(Video REC)旨在根据自然语言查询,在未剪辑的第一人称视频中定位被提及对象的时间与空间范围。然而,现有第一人称视频REC基准多聚焦短片段,目标密集出现,难以反映真实长时、未剪辑的视频特征。为此,本文提出LongEgoRefer,一个基于Ego4D数据集中长时视频构建的新基准。LongEgoRefer包含1,498个指代表达,平均视频时长达45分钟,具有极端目标稀疏性、详细语言描述及复杂的动态人-物交互。该任务要求模型在长时间序列中同时识别事件发生时间与目标位置,构成极具挑战性的时空定位问题。我们评估了包括基于视觉-语言模型的免训练基线在内的现有方法,结果表明即使先进模型也显著表现不佳,揭示了长时第一人称视频理解的内在难度,亟需更鲁棒的视频理解模型。

原文摘要 · Abstract (English)

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehension (Video REC), the task of localizing the temporal and spatial extent of a referred object in video frames given a natural language query, plays a key role in linking textual descriptions to observed objects in untrimmed egocentric recordings. However, existing egocentric Video REC benchmarks primarily focus on short video clips, where some target object appears densely within frames. Such settings do not reflect real-world egocentric recordings, which are long-form, untrimmed, and characterized by sparse object occurrences and complex activity transitions. To address this limitation, we introduce LongEgoRefer, a novel and challenging benchmark constructed from long-form videos in the Ego4D dataset. LongEgoRefer contains 1,498 referring expressions with an average video duration of 45 minutes. The benchmark exhibits extreme target sparsity, detailed linguistic descriptions, and complex human-object interactions embedded in long, dynamic egocentric narratives. Consequently, it defines a demanding spatio-temporal grounding problem that requires models to identify both when an event occurs and where the referred object appears within extended video sequences. We evaluate existing Video REC approaches, including training-free baselines based on vision-language models combined with Grounded SAM2. Extensive experiments show that even advanced baselines and current state-of-the-art models struggle significantly on LongEgoRefer. These results highlight the intrinsic difficulty of long-form egocentric spatio-temporal grounding and emphasize the need for more robust video understanding models.

视频理解指代消解长视频第一人称

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。