提出LongSpace框架,让模型能记住并回忆长视频中的空间信息。
LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

- 将长视频分块处理,融合3D结构信息和分层记忆
- 在空间推理任务上显著提升长视频理解能力
- 适合自动驾驶、机器人导航等需长期空间记忆的场景
多模态大语言模型(MLLMs)在图像与视频理解方面取得进展,能够处理更长的视觉输入。长时程任务如自动驾驶和机器人导航不仅需要识别当前视图,还需记住并回溯先前观察到的空间布局、路径、视角变化和物体状态。为此,我们提出了LongSpace-Bench——一个用于评估长时程空间记忆的房间巡游视频基准,涵盖场景感知、空间关系与空间记忆。本文进一步提出LongSpace,一种面向长视频空间推理的记忆框架。该框架将长视频建模为连续片段,将3D结构线索融入早期解码器层,并构建面向问题引导的分层记忆。在多个空间推理基准上的实验表明,LongSpace显著提升了长视频空间理解能力,进一步证明显式空间记忆是长时程视频MLLM的关键能力。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving and robotic navigation require more than recognizing the current view, as models must remember and retrieve previously observed spatial layouts, routes, viewpoint changes, and object states. To evaluate this capability, we introduce LongSpace-Bench, a room-tour video benchmark for long-horizon spatial memory, covering scene perception, spatial relations, and spatial memory. In this work, we further propose LongSpace, a memory framework for long-video spatial reasoning. LongSpace models long videos as sequential chunks, incorporates 3D structural cues into early decoder layers, and constructs layer-aware memory for question-guided retrieval. Experiments on multiple spatial reasoning benchmarks show that LongSpace improves long-video spatial understanding, further demonstrating explicit spatial memory as a key capability for long-horizon video MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。