arXiv:2606.17183cs.RO2026-06

融合知识图谱与上下文记忆,提升长时自指导航视频问答准确率。

VL-MemKnG: Hybrid Memory with a Spatio-Temporal Knowledge Graph for Question Answering over Long Egocentric Navigation Trajectories

论文配图:VL-MemKnG: Hybrid Memory with a Spatio-Temporal Knowledge Graph for Question Answering over Long Egocentric Navigation Trajectories
图 1 · 摘自论文原文
  • 用时空知识图谱+分段上下文记忆,双轨存储视频证据。
  • 在长时推理任务中召回率提升至40.55%,比基线高9个百分点。
  • 适合需要跨时段证据整合的导航类视频问答研究者。

在长时自指导航视频中回答路径相关问题,需从遥远时间片段中检索并组织证据,同时保持空间与上下文一致性。尽管长上下文视觉-语言模型可实现高质量答案,但对长轨迹计算开销大且重复查询效率低。现有图结构方法如VL-KnG通过持久化时空知识图缓解此问题,但仅依赖图检索可能忽略更广泛的时序连续性与上下文线索。本文提出VL-MemKnG,一种混合记忆框架,在VL-KnG基础上结合时空知识图谱与持久化分段上下文记忆:知识图谱捕捉结构化关系与长程物体关联,分段记忆则保留全局时序上下文以支持长时程证据检索。一个联合检索与推理模块协同作用于两种记忆表示,生成基于证据的答案及时间有序的支持证据。我们还引入WalkieKnowledgeT+,作为面向长时导航视频问答的新基准,包含需跨非共现时刻聚合证据的时序分布推理任务。在WalkieKnowledgeT+上,VL-MemKnG将Top-1检索准确率从58%提升至67%,Recall@1从34.50%提升至40.55%,优于所有对比方法(包括Gemini 2.5 Pro和Qwen 3.5+),尤其在时序全局与分散聚合类问题上优势显著,证明了结构化关系记忆与分段上下文记忆融合的有效性,同时保持高效查询推理。

原文摘要 · Abstract (English)

Answering navigation-relevant questions over long egocentric videos requires retrieving and organizing evidence distributed across distant temporal moments while maintaining spatial and contextual consistency. Although long-context vision--language models can achieve strong answer quality, they are computationally expensive for long trajectories and inefficient for repeated querying. Recent graph-based approaches such as VL-KnG address this challenge through persistent spatio-temporal knowledge graphs, but graph-centric retrieval alone may underrepresent broader temporal continuity and contextual cues. We present VL-MemKnG, a hybrid memory framework that extends VL-KnG by combining a spatio-temporal knowledge graph with persistent segment-level contextual memory. The knowledge graph captures structured relational information and long-range object associations, while segment-level memory preserves broader temporal context for long-horizon evidence retrieval. A hybrid retrieval-and-reasoning module jointly operates over both memory representations to produce evidence-grounded answers and temporally organized supporting evidence. We also introduce WalkieKnowledgeT+, an extension of WalkieKnowledge for long-horizon navigation-oriented video question answering. The benchmark includes temporally distributed reasoning tasks requiring evidence aggregation across multiple non-cooccurring moments. On WalkieKnowledgeT+, VL-MemKnG improves Top-1 retrieval accuracy from 58% to 67% and Recall@1 from 34.50% to 40.55%, outperforming all compared methods, including Gemini 2.5 Pro and Qwen 3.5+. The gains are particularly pronounced on temporal-global and temporally scattered aggregation questions, demonstrating the benefits of combining structured relational memory with segment-level contextual memory while maintaining efficient query-time inference.

视频问答时空记忆知识图谱导航理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。