arXiv:2609.08385cs.CVcs.RO2026-09

用重力对齐线框替代深度图,实现更鲁棒的单目室内定位。

GALoc: Gravity Aligned Wireframes for Depth-Free Monocular Floorplan Localization

论文配图:GALoc: Gravity Aligned Wireframes for Depth-Free Monocular Floorplan Localization
图 1 · 摘自论文原文
  • 基于重力对齐线框构建几何约束,无需依赖脆弱的深度预测
  • 在Gibson数据集上100步序列定位成功率88%(基线68%)
  • 适合结构清晰但无深度信息的复杂室内场景,可主动放弃不可靠判断

平面图是紧凑且与外观无关的理想室内地图,但现有方法依赖在杂乱场景中易失效的深度网络。本文提出GALoc,一种以几何为核心的框架,将深度预测替换为通过构造满足垂直性和共面性的重力对齐线框。给定单目RGB图像、相机内参、相对位姿和IMU方向,GALoc构建编码垂直性与共面性的线性约束矩阵,并通过全局搜索找到使最小奇异值最小的相机姿态。校正后的线框通过闭式、视场一致的变换投影至鸟瞰图布局,并通过无度量的SE(2)搜索匹配平面图。我们在Structured3D上进行端到端评估,在带校准噪声的Gibson数据集及作者自采的真实序列上验证。当可见墙结构充足时,GALoc表现匹配或超越基于深度的基线:在Gibson上100步序列0.1m精度下定位成功率达88%(基线68%),并在结构缺失场景中主动放弃判断。

原文摘要 · Abstract (English)

Floorplans are compact, appearance-invariant maps ideal for indoor localization, yet existing methods rely on depth networks that are brittle in cluttered scenes. We propose GALoc, a geometry-first framework that replaces depth prediction with gravity-aligned wireframes that satisfy verticality and coplanarity by construction. Given monocular RGB, camera intrinsics, relative poses, and IMU orientation, GALoc constructs a linear constraint matrix encoding verticality and coplanarity, and finds the camera gauge minimizing its smallest singular value via global search. The rectified wireframes are projected into bird's-eye-view layouts through a closed-form, FOV-consistent transformation and matched against the floorplan via metric-free SE(2) search. We evaluate end-to-end on Structured3D, with calibrated noise on Gibson, and on real-world author-collected sequences. When sufficient wall geometry is visible, GALoc matches or outperforms depth-based baselines -- achieving 88% sequential localization success at 0.1m over 100-step sequences on Gibson vs the baseline's 68% -- while abstaining in structure-blind scenes.

单目定位几何先验平面图匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。