发现视觉模型隐空间自发具备3D欧氏几何特性,可直接用于导航定位。
SeeSE3: Emergence of 3D Space in Vision Features

- 通过拓扑与几何双重探针,检验特征空间与SE(3)变换群的关联性。
- 自监督模型隐空间与3D欧氏空间相关性极强,无需显式3D训练。
- 提出隐空间导航新方法,跳过3D重建直接实现视觉里程计与定位。
本文探讨视觉基础模型是否构建出反映三维欧几里得空间内在性质的表征。不同于以往通过回归图像中心量(如深度或法向)来探测视觉特征的3D感知,我们从拓扑和几何两个角度考察视觉特征空间结构与欧氏变换群SE(3)的关系。提出一套探针:基于互邻域度量的拓扑对齐测试,以及用于检验相机运动几何线性可达性的Poincaré Adapter。结果表明,尽管自监督视觉模型未接受直接3D监督或主动交互训练,其隐空间仍与三维欧氏空间表现出显著强相关性。基于此洞察,我们提出一类新的“隐空间导航”技术,可在不进行显式3D重建的前提下,仅在隐空间内完成视觉里程计与定位任务。
原文摘要 · Abstract (English)
In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that probe 3D awareness of vision features by regressing image-centric quantities such as depth or normals, we investigate the relation between the structure of the space of visual features and the group of Euclidean transformations $SE(3)$. We propose a set of probes to evaluate this relation from both topological and geometric perspectives: a mutual neighborhood metric that measures the alignment between feature neighborhoods and spatial topology, and a Poincaré Adapter to test the linear accessibility of the geometry of camera motion from latent displacements in static scenes. We show that self-supervised vision models, which, in principle, have not been trained with direct 3D supervision or active agency, possess latent subspaces that are remarkably strongly correlated with three-dimensional Euclidean space, when probed correctly. Building on this insight we propose a new class of "Latent-Space Navigation" techniques that perform visual odometry and localization purely in the latent space, bypassing the need for explicit 3D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。