arXiv:2601.14895cs.CVcs.AI2026-01被引 1

用三维空间结构做记忆索引,让模型能理解长时间视频中的位置关系。

SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA

  • 基于3D结构构建可解释的视频记忆骨架,支持空间定位与检索
  • 在真实室内场景中实现稳定的位置推理与路径引导,抗遮挡能力强
  • 适合需要长期视觉理解的任务,如语言问答和导航辅助

我们提出SpatialMem,一种面向自视角长时视频的语言对齐检索与问答的记忆中心系统。该系统以度量3D空间为可解释的索引框架,而非显式映射目标。从随意拍摄的自视角RGB视频出发,SpatialMem构建室内场景的度量对齐空间结构,检测墙壁、门、窗等3D结构锚点作为第一层支撑,并在层次化记忆中填充开放词汇物体节点,将证据片段、视觉嵌入及双层文本描述关联到3D坐标,实现紧凑存储与快速检索。该设计支持对距离、方向、可见性等空间关系的可解释查询,适用于语言引导检索/问答和预建记忆下的离线导航引导,无需专用传感器。在公开的Replica场景及两个真实自视角室内场景上的实验表明,SpatialMem在复杂遮挡和杂乱环境下仍保持稳定的布局推理、离线引导与层次化检索能力。消融实验显示,双层文本描述提升路径级语义对齐,适度规模扰动仅导致有限性能下降。结果表明,SpatialMem是空间对齐长时视频理解的一种高效可扩展记忆接口。

原文摘要 · Abstract (English)

We present SpatialMem, a memory-centric system for long-horizon, language-grounded retrieval and QA from egocentric video, where metric 3D serves as an interpretable indexing scaffold rather than an explicit mapping objective. Starting from casually captured egocentric RGB video, SpatialMem builds a metric-aligned spatial scaffold for indoor scenes, detects structural 3D anchors (walls, doors, windows) as first-layer support, and populates a hierarchical memory with open-vocabulary object nodes that link evidence patches, visual embeddings, and two-layer textual descriptions to 3D coordinates for compact storage and fast retrieval. This design enables interpretable, spatially grounded queries over relations (e.g., distance, direction, visibility) and supports downstream tasks such as language-guided retrieval/QA and offline navigation-style guidance over a prebuilt memory, without specialized sensors. Experiments on one public Replica scene and two real-world egocentric indoor scenes show that SpatialMem maintains stable layout reasoning, offline guidance, and hierarchical retrieval across these evaluated scenes despite increasing clutter and occlusion. A compact ablation further shows that the two-layer description memory improves path-level grounding, while moderate scale perturbation causes only limited degradation. These results position SpatialMem as an efficient and extensible memory interface for spatially grounded long-horizon video understanding.

视频记忆空间推理语言对齐自视角视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。