arXiv:2606.15287cs.CV2026-06

通过几何引导与实例感知,提升图像到点云的场景识别精度。

G2IA: Geometry-Guided Instance-Aware Retrieval and Refinement for Cross-Modal Place Recognition

论文配图:G2IA: Geometry-Guided Instance-Aware Retrieval and Refinement for Cross-Modal Place Recognition
图 1 · 摘自论文原文
  • 融合视觉几何先验与实例特征构建匹配度更高的场景描述符。
  • 在召回结果中通过局部形状与空间布局一致性重排序,提升准确性。
  • 适用于城市环境下的机器人定位,跨数据集泛化能力强。

跨模态场景识别(CMPR)使仅依赖摄像头的机器人能在自主导航中定位至预先构建的激光雷达地图。该图像到点云的任务面临双重挑战:透视RGB外观与稀疏度量几何之间的模态差异,以及城市场景中道路、立面、交叉口和物体布局相似导致的感知混淆。不同于将CMPR视为单一全局描述符匹配问题,本文认为可靠检索需同时实现几何感知表示对齐与细粒度候选验证。为此提出G2IA框架:在检索阶段,融合VGGT提供的视觉几何先验与实例特征,生成更契合激光雷达地图表示的场景描述符;在精炼阶段,通过显式验证局部实例形状及其相对空间布局在不同模态间的一致性,对召回候选进行重排序。在多个公开基准上的实验表明,G2IA在不同定位阈值下均显著提升图像到点云的场景识别性能,并展现出强跨数据集泛化能力。

原文摘要 · Abstract (English)

Cross-modal place recognition (CMPR) enables camera-only robots to localize against pre-built LiDAR maps in autonomous navigation scenarios. This image-to-point-cloud setting is challenged by two coupled ambiguities: the modality gap between perspective RGB appearance and sparse metric geometry, and perceptual aliasing among urban places with similar roads, facades, intersections, and object arrangements. Instead of treating CMPR as a single global descriptor matching problem, we argue that reliable retrieval requires both geometry-aware representation alignment and fine-grained candidate verification. In this paper, we propose G2IA, a geometry-guided instance-aware framework for image-to-point-cloud place recognition. In the retrieval stage, visual geometry priors from VGGT and instance features are integrated to construct place descriptors that are more compatible with LiDAR-derived map representations. In the refinement stage, the retrieved candidates are re-ranked by explicitly verifying whether local instance shapes and their relative spatial layouts are consistent across modalities. Experiments on public benchmarks demonstrate that G2IA consistently improves image-to-point-cloud place recognition under different localization thresholds, and exhibits strong cross-dataset generalization.

场景识别跨模态几何先验机器人定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。