arXiv:2605.09449cs.CV2026-05

让视频多模态模型学会构建3D空间认知地图,实现跨视角空间一致性推理。

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs

论文配图:SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
图 1 · 摘自论文原文
  • 基于哺乳动物双流系统,用体素化地图整合视觉信息
  • 在VSI-Bench上达到新基准,跨场景泛化能力显著提升
  • 适合需要3D空间理解的机器人、自动驾驶等应用

近期多模态大模型在视觉理解和语言推理方面取得显著进展,但在三维环境中缺乏持久的世界中心表示,难以实现空间一致推理。受哺乳动物双流系统的启发,我们提出SpaceMind++,一种视频多模态大模型架构,从RGB视频中显式构建体素化的认知地图。该地图将碎片化的自我中心观测重组为共享的三维度量表示,使模型能够保持物体恒存性和空间拓扑结构,即使视角变化也能维持一致性。为使这一世界中心表示可被预训练视频多模态大模型使用而不破坏其原有视觉标记接口,我们引入坐标引导的深度迭代融合机制,通过坐标嵌入和3D旋转位置编码,将地图级空间知识回传至原始二维视觉特征,实现语义交互在度量三维空间中的定位,类似内嗅皮层对感觉特征的度量空间绑定。大量实验表明,SpaceMind++在VSI-Bench上达到新最优性能,并在SPBench、SITE-Bench和SPAR-Bench上展现出优越的分布外泛化能力,证明其在未见三维环境中的鲁棒性。

原文摘要 · Abstract (English)

Recent multimodal large language models (MLLMs) have made remarkable progress in visual understanding and language-based reasoning, yet they lack a persistent world-centered representation for spatially consistent reasoning in 3D environments. Inspired by the mammalian dual-stream system, where semantic and spatial cues are processed separately and integrated into an allocentric cognitive map, we propose SpaceMind++, a video MLLM architecture that explicitly builds a voxelized cognitive map from RGB videos. This map reorganizes fragmented egocentric observations into a shared 3D metric representation, enabling the model to preserve object permanence and spatial topology across changing viewpoints. To make this allocentric representation usable by a pretrained video MLLM without disrupting its native visual-token interface, we introduce Coordinate-Guided Deep Iterative Fusion, a new mechanism that relays map-level spatial knowledge back into the original 2D visual features. This fusion is explicitly guided by coordinate embeddings and 3D Rotary Positional Encoding, which ground semantic interactions in metric 3D space, resembling the entorhinal binding of sensory features to metric space. Extensive experiments show that SpaceMind++ achieves new state-of-the-art performance on VSI-Bench. Furthermore, it demonstrates superior out-of-distribution generalization on SPBench, SITE-Bench, and SPAR-Bench, underscoring its robustness in unseen 3D environments.

空间认知多模态模型3D理解视频推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。