无需语义标注,通过几何对比消除重复布局导致的定位混淆。
DisCo-FLoc: Semantic-Free Floorplan Localization via $SE(2)$-Aware Contrastive Disambiguation
- 用深度感知的射线回归预测器将图像转为2D射线,生成几何感知候选位姿。
- 通过SE(2)扰动构建正负样本,实现位置与方向的精准区分。
- 在无语义标签下超越现有方法,尤其提升方向定位精度。
视觉楼层平面定位(FLoc)面临严重结构混淆问题,源于重复的极简布局导致物理上相距较远的位姿具有高度相似的视觉-几何特征,从而降低空间可分性和角度可辨性。现有方法依赖昂贵的语义标注来缓解此问题,但性能提升有限。为此,本文提出DisCo-FLoc,一种无需语义信息的视觉-几何对比消歧方法。首先,引入深度感知的射线回归预测器(RRP),作为从密集图像到射线的几何投影器,通过抑制垂直方向视觉杂波,将单目RGB图像映射为2D射线基元,并与楼图匹配生成几何感知的定位候选。其次,为解决候选间的剩余歧义,提出空间扰动的对比目标,对齐RGB图像与局部楼图结构,并构建视觉-几何兼容函数。特别地,通过SE(2)位姿扰动在位置和方向层面精细构造正负样本进行对比学习,有效实现位姿平滑性、空间可分性和角度可辨性。该兼容函数使DisCo-FLoc能利用更丰富的视觉上下文,而非仅依赖几何布局进行定位消歧,且无需任何语义标注。在两个挑战性视觉FLoc基准上的大量实验表明,DisCo-FLoc显著优于当前最优的语义依赖方法,尤其缩小了位置与方向定位精度之间的差距。
原文摘要 · Abstract (English)
Visual Floorplan Localization (FLoc) struggles with severe structural aliasing caused by repetitive minimalist layouts. This occurs because physically distant poses share highly similar visual-geometric features, which degrades spatial separability and angular discriminability. While existing methods attempt to mitigate these ambiguities by relying on costly semantic annotations, the resulting performance gains remain inherently limited. To address the above issues, we propose DisCo-FLoc, a semantic-free method for visual-geometric Contrastive Disambiguation. First, we introduce a depth-aware Ray Regression Predictor (RRP) that serves as a dense-to-ray geometric projector. By explicitly suppressing visual clutter along the vertical dimension, RRP projects monocular RGB images into 2D ray primitives, which are matched with floorplans to produce geometry-aware FLoc candidates. Second, to resolve the remaining ambiguity among these candidates, we propose a spatially perturbed contrastive objective to align RGB images with local floorplan structures and formulate a visual-geometric compatibility function. In particular, we meticulously construct positive and negative samples at both positional and directional levels through $SE(2)$ pose perturbations for contrastive learning, effectively achieving pose smoothness, spatial separability, and angular discriminability. The compatibility function enables DisCo-FLoc to disambiguate FLoc by using richer visual context beyond pure geometric layouts, without requiring any semantic annotations. Extensive experiments on two challenging visual FLoc benchmarks demonstrate that DisCo-FLoc significantly outperforms state-of-the-art semantic-based methods, especially narrowing the performance gap between positional and directional FLoc accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。