让长时第一视角视频可主动查询,提升理解能力
AMEGO: Active Memory from long EGOcentric videos

- 从单段长视频构建自包含语义无关表征,支持多轮查询
- 在2万+挑战性查询上超越基线模型,显著提升视频理解性能
- 适合研究长视频推理、记忆建模与人机交互的开发者
第一视角视频提供了个人日常经历的独特视角,但其非结构化特性给感知带来挑战。本文提出AMEGO,一种新方法,旨在增强对超长第一视角视频的理解。受人类单次观看即能保持信息能力的启发,AMEGO专注于从一段第一视角视频中构建自包含表征,捕捉关键位置与物体交互。该表征为语义无关,支持无需重新处理全部视觉内容即可进行多轮查询。此外,为评估对超长第一视角视频的理解能力,我们引入新的主动记忆基准(Active Memories Benchmark, AMB),包含来自EPIC-KITCHENS的超过20,000个高难度视觉查询。这些查询涵盖不同层次的视频推理任务(顺序性、并发性与时间定位),用于评估细粒度视频理解能力。我们在AMB上展示了AMEGO的优异表现,显著优于其他视频问答基线。
原文摘要 · Abstract (English)
Egocentric videos provide a unique perspective into individuals' daily experiences, yet their unstructured nature presents challenges for perception. In this paper, we introduce AMEGO, a novel approach aimed at enhancing the comprehension of very-long egocentric videos. Inspired by the human's ability to maintain information from a single watching, AMEGO focuses on constructing a self-contained representations from one egocentric video, capturing key locations and object interactions. This representation is semantic-free and facilitates multiple queries without the need to reprocess the entire visual content. Additionally, to evaluate our understanding of very-long egocentric videos, we introduce the new Active Memories Benchmark (AMB), composed of more than 20K of highly challenging visual queries from EPIC-KITCHENS. These queries cover different levels of video reasoning (sequencing, concurrency and temporal grounding) to assess detailed video understanding capabilities. We showcase improved performance of AMEGO on AMB, surpassing other video QA baselines by a substantial margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。