arXiv:2507.18881cs.CVcs.RO2025-07中稿 · ACM MM 2025被引 4

用3D几何先验提升视觉楼层图定位精度,无需额外标注

Perspective from a Higher Dimension: Can 3D Geometric Priors Help Visual Floorplan Localization?

  • 通过多视角几何约束建模3D视角不变性
  • 结合场景表面重建增强视觉与平面的对应关系
  • 自监督学习实现无标注3D先验,提升定位准确率

由于建筑楼层图易于获取、长期一致且对视觉变化鲁棒,基于楼层图的自定位受到关注。但楼层图作为结构的极简表示,其与视觉感知在模态和几何上的差异带来了挑战。现有方法虽利用2D几何特征和姿态滤波取得良好效果,仍难以应对因3D物体形状导致的频繁视觉变化和视图遮挡。本文从高维视角出发,将3D几何先验注入视觉楼层图定位算法中。首先通过多视角几何约束建模几何感知的视角不变性;其次构建视图-场景对齐的几何先验,通过关联场景表面重建与序列RGB帧增强跨模态几何-颜色对应。两种3D先验均采用自监督对比学习建模,无需额外几何或语义标注。在大量真实场景中的实验证明,该方法显著优于当前最优方法,大幅提升定位成功率,且未增加原定位算法的计算开销。所有数据与代码将在匿名评审后公开。

原文摘要 · Abstract (English)

Since a building's floorplans are easily accessible, consistent over time, and inherently robust to changes in visual appearance, self-localization within the floorplan has attracted researchers' interest. However, since floorplans are minimalist representations of a building's structure, modal and geometric differences between visual perceptions and floorplans pose challenges to this task. While existing methods cleverly utilize 2D geometric features and pose filters to achieve promising performance, they fail to address the localization errors caused by frequent visual changes and view occlusions due to variously shaped 3D objects. To tackle these issues, this paper views the 2D Floorplan Localization (FLoc) problem from a higher dimension by injecting 3D geometric priors into the visual FLoc algorithm. For the 3D geometric prior modeling, we first model geometrically aware view invariance using multi-view constraints, i.e., leveraging imaging geometric principles to provide matching constraints between multiple images that see the same points. Then, we further model the view-scene aligned geometric priors, enhancing the cross-modal geometry-color correspondences by associating the scene's surface reconstruction with the RGB frames of the sequence. Both 3D priors are modeled through self-supervised contrastive learning, thus no additional geometric or semantic annotations are required. These 3D priors summarized in extensive realistic scenes bridge the modal gap while improving localization success without increasing the computational burden on the FLoc algorithm. Sufficient comparative studies demonstrate that our method significantly outperforms state-of-the-art methods and substantially boosts the FLoc accuracy. All data and code will be released after the anonymous review.

视觉定位3D先验自监督学习楼层图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。