arXiv:2605.27318cs.CV2026-05

用问题引导的几何记忆提升视频空间推理能力

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning

论文配图:Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
图 1 · 摘自论文原文
  • 根据问题动态选择几何证据,避免无关信息干扰
  • 在多个基准上达到当前最优性能,长时推理更准确
  • 适合需要精准空间理解的视频问答任务

视频空间推理需随时间积累视角相关的证据,同时保留与问题相关的信息。现有视频多模态模型虽提升了几何感知和长程上下文建模能力,但常将记忆视为通用时间缓存,易引入冗余或无关证据,削弱长时推理效果。本文提出Q-GeoMem,一种面向视频空间推理的问题引导几何记忆框架。该框架将相机条件下的几何信息注入视觉标记,并维护两个互补记忆:用于近期密集特征与相机状态的细粒度上下文库,以及用于紧凑长程证据的语义-几何证据库。对每个候选帧,通过校准的Q-Former估计问题相关性,同时基于活跃证据库重新计算新颖性和证据效用。所得相关性-新颖性效用控制容量替换策略,并作为注意力偏置用于记忆读取。推理时,两个记忆在更新前被读取,并与当前帧表示自适应融合。大量实验在两个域内及五个分布外基准上进行,结合受控记忆分析表明,Q-GeoMem在评估设置中达到最先进水平,验证了问题引导几何证据选择的有效性。

原文摘要 · Abstract (English)

Video spatial reasoning requires accumulating viewpoint-dependent evidence over time while retaining information useful to the question being asked. Existing spatial video-language models improve geometric perception and long-range context modeling, but often treat memory as a generic temporal cache, which can introduce redundant or irrelevant evidence and weaken long-horizon reasoning. We propose Q-GeoMem, a question-guided geometric memory framework for video spatial reasoning. Q-GeoMem injects camera-conditioned geometry into visual tokens and maintains two complementary memories: a Fine-Grained Context Bank for recent dense features and camera states, and a Semantic-Geometric Evidence Bank for compact long-range evidence. For each candidate frame, a calibrated Q-Former estimates question relevance, while novelty and evidence utility are recomputed with respect to the active evidence bank. The resulting relevance-novelty utility controls capacity-based replacement and serves as an attention bias during memory reading. During reasoning, both memories are read before update and adaptively fused with the current frame representation. Extensive experiments across two in-domain and five out-of-distribution benchmarks, and controlled memory analyses show that Q-GeoMem achieves state-of-the-art performance in the evaluated settings and validate the effectiveness of question-guided geometric evidence selection.

视频推理几何记忆多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。