arXiv:2602.23709cs.CV2026-02被引 2

构建动态知识图谱,实现多日第一视角视频的长期时序推理

EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding

  • 基于第一视角场景设计新型图结构,统一建模人、物、地点等实体
  • 通过时序关系建模,在多日视频中累积稳定长时记忆
  • 无需训练,适用于长时视频问答,适合研究长期视觉理解的学者

跨越多日的超长第一视角视频给视频理解带来巨大挑战。现有方法仍依赖片段化局部处理和有限时序建模,难以对长时间序列进行推理。为此,我们提出EgoGraph,一种无需训练的动态知识图构建框架,显式编码第一视角视频流中的长期跨实体依赖关系。EgoGraph采用新颖的第一视角模式,统一提取与抽象人物、物体、地点、事件等核心实体,并结构化地推理其属性与交互,生成比传统基于片段的视频模型更丰富、更连贯的语义表征。关键在于,我们设计了时序关系建模策略,捕捉实体间的跨时间依赖,并在多日数据上累积稳定的长期记忆,支持复杂时序推理。在EgoLifeQA与EgoR1-bench基准上的大量实验表明,EgoGraph在长期视频问答任务上达到领先性能,验证了其作为超长第一视角视频理解新范式的有效性。

原文摘要 · Abstract (English)

Ultra-long egocentric videos spanning multiple days present significant challenges for video understanding. Existing approaches still rely on fragmented local processing and limited temporal modeling, restricting their ability to reason over such extended sequences. To address these limitations, we introduce EgoGraph, a training-free and dynamic knowledge-graph construction framework that explicitly encodes long-term, cross-entity dependencies in egocentric video streams. EgoGraph employs a novel egocentric schema that unifies the extraction and abstraction of core entities, such as people, objects, locations, and events, and structurally reasons about their attributes and interactions, yielding a significantly richer and more coherent semantic representation than traditional clip-based video models. Crucially, we develop a temporal relational modeling strategy that captures temporal dependencies across entities and accumulates stable long-term memory over multiple days, enabling complex temporal reasoning. Extensive experiments on the EgoLifeQA and EgoR1-bench benchmarks demonstrate that EgoGraph achieves state-of-the-art performance on long-term video question answering, validating its effectiveness as a new paradigm for ultra-long egocentric video understanding.

第一视角视频知识图谱时序推理长时记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。