首个面向第一人称视频的像素级时空定位基准,解决小目标、短时距等挑战。
Fine-grained Spatiotemporal Grounding on Egocentric Videos
- 构建自动标注流水线,生成第一人称视频中物体掩码与指代表达
- 在短/中/长视频上标注了大量实例,验证主流模型性能下降超30%
- 适合作为第一人称视觉理解研究者的基准工具
时空视频定位旨在根据文本查询定位视频中的目标实体。尽管现有研究在第三人称视频上取得显著进展,但第一人称视频场景仍相对未被充分探索,而其在增强现实和机器人等应用中日益重要。本文系统分析了第一人称与第三人称视频的差异,揭示关键挑战:目标持续时间更短、轨迹更稀疏、目标尺寸更小、位置偏移更大。为此,我们提出EgoMask,首个针对第一人称视频的像素级细粒度时空定位基准。该数据集通过我们提出的自动标注流程构建,涵盖短、中、长期视频中的指代表达与物体掩码。同时,我们创建了EgoMask-Train大规模训练数据集以促进模型发展。实验表明,现有先进模型在EgoMask上的表现显著下降,但在EgoMask-Train上微调后性能大幅提升,且保持在第三人称数据集上的表现。本工作为推进第一人称视频理解提供了关键资源与洞见。代码已开源。
原文摘要 · Abstract (English)
Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite its growing importance in applications such as augmented reality and robotics. In this work, we conduct a systematic analysis of the discrepancies between egocentric and exocentric videos, revealing key challenges such as shorter object durations, sparser trajectories, smaller object sizes, and larger positional shifts. To address these challenges, we introduce EgoMask, the first pixel-level benchmark for fine-grained spatiotemporal grounding in egocentric videos. It is constructed by our proposed automatic annotation pipeline, which annotates referring expressions and object masks across short-, medium-, and long-term videos. Additionally, we create EgoMask-Train, a large-scale training dataset to facilitate model development. Experiments demonstrate that the state-of-the-art spatiotemporal grounding models perform poorly on our benchmark EgoMask, but fine-tuning on EgoMask-Train yields significant improvements, while preserving performance on exocentric datasets. Our work thus provides essential resources and insights for advancing egocentric video understanding. Our code is available at https://github.com/LaVi-Lab/EgoMask .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。