arXiv:2511.20644cs.CV2025-11中稿 · ECCV被引 4

用双记忆机制提升视频空间推理能力,让模型像人一样理解3D场景。

Vision-Language Memory for Spatial Reasoning

  • 设计工作记忆与情景记忆双模块,持续保留跨帧空间信息。
  • 在多个基准上达到当前最佳性能,显著提升视频空间推理准确率。
  • 适合需要长期视觉理解的机器人导航、自动驾驶等应用。

空间推理是智能机器人的重要能力,但现有视觉语言模型在基于视频的空间推理上仍远未达到人类水平。主要瓶颈在于语义与几何信息不一致,以及缺乏持续记忆以保留跨帧的3D表征。为此,我们提出VLM$^2$,一种具有持久记忆的视觉语言模型,仅从2D视频中构建视图一致且具备3D感知能力的表示。其核心是双记忆模块:工作记忆作为滑动窗口聚焦当前上下文,情景记忆则跨帧整合并存储关键信息。该设计在固定计算开销下实现高效且有界的时空推理。在多个基准上的大量实验表明,VLM$^2$在视频类模型中达到领先性能,显著推进了视觉空间智能的边界。

原文摘要 · Abstract (English)

Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges: a semantic-geometric misalignment that prevents consistent 3D understanding, and the absence of persistent memory to retain 3D representation and understanding across frames. To address these limitations, we present VLM$^2$, a Vision-Language Model with persistent Memory for spatial reasoning with a view-consistent, 3D-aware representation purely from 2D videos. Specifically, we incorporate a dual-memory module consisting of a working memory that operates as a sliding window to focus on immediate context, and an episodic memory that consolidates and stores critical information across frames. This design enables bounded and efficient spatial reasoning under a fixed computational cost. Extensive experiments on multiple benchmarks show that VLM$^2$ achieves state-of-the-art performance among video-based models, significantly advancing the frontier of visual-spatial intelligence.

空间推理视觉语言模型记忆机制视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。