用场景图压缩长视频,让大模型能完整理解第一人称视频。
Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs

- 将视频转为带时间戳的结构化场景图,大幅减少输入长度。
- 在HD-EPIC数据集上超越现有基线,实现最先进性能。
- 适合需要长视频推理的智能助手、自动驾驶等场景。
现有多模态大模型因输入令牌限制,在处理长视频序列时面临挑战。尤其在第一人称视频中,由于动态复杂、状态频繁变化和摄像头移动,当前方法被迫大量帧采样,导致严重的时间与上下文信息丢失,制约细粒度视频推理能力。本文提出一种第一人称视频问答框架,通过第一人称场景图(EgoSGs)克服这一限制——即带有时间锚定的结构化表示,捕捉物体、属性、空间关系及交互随时间的变化。将视频转化为紧凑的文本型场景图,以符号形式保留原始视频的核心视觉与时间信息,显著降低输入长度同时保持语义丰富性。关键在于,该表示使大模型能在其令牌预算内高效推理整个视频序列。在HD-EPIC VQA数据集上,本方法取得领先结果,优于多个强基线模型,表明如EgoSGs这类结构化、时间锚定的表示可弥合长视频理解与当前大模型上下文限制之间的鸿沟。
原文摘要 · Abstract (English)
Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations. As a result, current video understanding approaches, especially in egocentric settings characterized by complex dynamics, frequent state changes, and moving cameras, are forced to massively subsample frames. This leads to severe loss of temporal and contextual information, constraining their ability to perform fine-grained video reasoning. In this work, we introduce a framework for egocentric video question answering (VQA) that overcomes these input constraints through Egocentric Scene Graphs (EgoSGs), i.e., temporally grounded, structured representations that capture objects, attributes, spatial relations, and interactions over time. By representing videos as compact, text-based scene graphs, our method preserves the essential visual and temporal information of the original video in a symbolic form that drastically reduces input length while maintaining semantic richness. Crucially, this enables MLLMs to reason efficiently over entire video sequences within their token budget. On HD-EPIC VQA, our method achieves state-of-the-art results, outperforming strong video-based baselines on multiple models and suggesting that structured, temporally grounded representations like EgoSGs can bridge long-form egocentric video understanding and the context limitations of today's MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。