arXiv:2502.11742cs.CV2025-02被引 1

融合距离图与俯视图,提升图像到点云的场景识别精度

Range and Bird's Eye View Fused Cross-Modal Visual Place Recognition

  • 提出初始检索+重排序框架,融合垂直与水平视角信息
  • 在KITTI数据集上超越现有方法,召回率显著提升
  • 适合需要高精度定位的自动驾驶场景

图像到点云的跨模态视觉场景识别是一项挑战性任务,查询为RGB图像,数据库样本为激光雷达点云。相比单模态方法,该方式结合了摄像头广泛部署与点云在空间几何和距离信息上的鲁棒性。然而,现有方法依赖于仅捕捉垂直或水平视场的中间模态,难以充分利用双传感器互补信息。本文提出一种创新的初始检索+重排序方法,有效融合距离图(range)与鸟瞰图(BEV)信息。方法仅通过高效的全局描述符相似性搜索实现重排序,并引入新颖的相似性标签监督机制,以最大化有限训练数据的利用率。具体地,使用点云平均距离近似外观相似性,并在基础三元组损失中引入基于相似性差异的自适应边界。在KITTI数据集上的实验表明,本方法显著优于当前最先进方法。

原文摘要 · Abstract (English)

Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are LiDAR point clouds. Compared to single-modal VPR, this approach benefits from the widespread availability of RGB cameras and the robustness of point clouds in providing accurate spatial geometry and distance information. However, current methods rely on intermediate modalities that capture either the vertical or horizontal field of view, limiting their ability to fully exploit the complementary information from both sensors. In this work, we propose an innovative initial retrieval + re-rank method that effectively combines information from range (or RGB) images and Bird's Eye View (BEV) images. Our approach relies solely on a computationally efficient global descriptor similarity search process to achieve re-ranking. Additionally, we introduce a novel similarity label supervision technique to maximize the utility of limited training data. Specifically, we employ points average distance to approximate appearance similarity and incorporate an adaptive margin, based on similarity differences, into the vanilla triplet loss. Experimental results on the KITTI dataset demonstrate that our method significantly outperforms state-of-the-art approaches.

视觉定位跨模态激光雷达自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。