arXiv:2607.17599cs.CV2026-07

让视频空间推理更稳定,靠几何一致性建记忆与强化学习

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

论文配图:ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
图 1 · 摘自论文原文
  • 用几何一致性构建记忆,整合冗余视频中的空间证据
  • 在三个基准上平均得分提升12.6分,超越最强基线
  • 适合需要跨视角稳定推理的导航与长视频问答任务

视频空间推理对导航感知和长视频问答至关重要,要求模型在视角变化下长期推断空间关系。然而现有多模态大模型仍以语义为中心,难以可靠聚合来自冗余视频观测的一致空间证据,导致推理效率低或不稳定。为此,我们提出ConsiSpace,一种面向几何敏感视频空间推理的几何一致性感知框架,将空间一致性作为证据组织原则和显式后SFT学习信号。构建包含隐式证据标记与显式几何线索的几何一致记忆(GCM),并采用高效组织策略紧凑保存任务相关空间证据。此外,在监督微调后使用统一一致性自监督强化学习(UC-SSRL)提升跨视图稳定性,引入答案、度量和拓扑一致性奖励。在三个空间推理基准VSI-Bench、OSI-Bench和MMSI-Video-Bench上的大量实验显示,性能持续提升,平均得分比最强基线高出12.6分。

原文摘要 · Abstract (English)

Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.

视频推理空间关系几何一致性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。