arXiv:2605.07148cs.CV2026-05被引 1

发现并操控视觉语言模型中的3D拓扑潜空间,提升空间推理能力。

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models

论文配图:Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models
图 1 · 摘自论文原文
  • 通过线性提取分离出模型的纯空间潜变量
  • 在真实场景任务中提升12.1%的空间理解性能
  • 适合研究视觉模型空间表征与可控生成的学者

几十年的认知科学研究表明,人类通过形成以环境为中心、保持拓扑结构的3D空间认知地图来导航。尽管现代视觉语言模型(VLMs)能从2D视角输入中涌现出空间推理能力,但其是否构建了类似的3D内部表示仍不明确。本文证明当前VLM确实具备3D场景的潜在拓扑地图,但被颜色、形状等非几何语义强烈掩盖。通过跨场景线性特征提取,我们分离出一个纯净的空间子空间,该子空间可因果控制模型的空间输出。我们数学上塑造这一潜表示,并证明其对应于场景3D高斯核图的拉普拉斯特征映射,在连续极限下收敛至真实3D空间。基于此几何识别,我们提出一种基于狄利克雷能量的数学严谨潜正则化方法。仅在500步监督微调(SFT)中引入该单项正则化,即在真实世界空间基准测试中显著提升表现,相较标准SFT和竞争基线最高提升12.1%,尤其在场景拓扑理解任务中效果突出。源代码已公开于https://github.com/pittisl/vlm-latent-shaping。

原文摘要 · Abstract (English)

Decades of cognitive science establish that humans navigate environments by forming cognitive maps, defined as allocentric and topology-preserving representations of 3D space. While modern Vision-Language Models (VLMs) demonstrate emergent spatial reasoning from 2D egocentric inputs, it remains unclear whether they construct an analogous 3D internal representation. In this paper, we demonstrate that current VLMs do possess a latent topological map of 3D scenes, but it is heavily overshadowed by non-geometric visual semantics, such as color and shape. By isolating this spatial subspace through cross-scene linear feature extraction, we extract a clean spatial subspace that causally controls the model's spatial outputs. We mathematically shape this latent representation and prove its correspondence to the Laplacian eigenmaps of the scene's 3D Gaussian-kernel graph, converging to the physical 3D space in the continuous limit. Motivated by this geometric identification, we further introduce a mathematically principled latent regularization method for VLMs, based on Dirichlet energy. Applying this single-term regularizer to a minimal 500-step supervised VLM fine-tuning (SFT) on simple synthetic data yields significant improvements on real-world spatial benchmarks, outperforming standard SFT and competitive baselines by up to 12.1\% in spatial tasks involving scene topology understanding. Source code is available at https://github.com/pittisl/vlm-latent-shaping

视觉语言模型3D拓扑潜空间空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。