arXiv:2412.15577cs.CVcs.LG2024-12被引 6

用对比学习融合图像与点云,提升无信号环境下的定位精度。

SaliencyI2PLoc: saliency-guided image-point cloud localization using contrastive learning

  • 通过显著性图引导特征聚合,增强全局特征表达。
  • 在城市场景下召回率@1达78.92%,较基线提升37.35%。
  • 适合多机器人地图融合与城市资产管理系统使用。

图像到点云的全局定位在无GNSS环境的机器人导航中至关重要,对多机器人地图融合与城市资产管理日益重要。图像与点云间的模态差距给跨模态融合带来挑战。现有方法或需模态统一(导致信息丢失),或依赖人工设计训练方案(缺乏特征对齐与关系一致性)。为此,本文提出SaliencyI2PLoc,一种基于对比学习的架构,将显著性图融入特征聚合,并在多流形空间中保持特征关系一致性。为减少数据预处理负担,采用对比学习框架实现高效跨模态特征映射。设计上下文显著性引导的局部特征聚合模块,充分挖掘场景中静态信息,生成更具代表性的全局特征。同时,在对比学习中考虑不同流形空间样本间相对关系的一致性,以增强跨模态特征对齐。在城市与高速公路场景数据集上的实验表明,本方法有效且鲁棒。具体地,在城市场景评估数据集上,Recall@1达到78.92%,Recall@20达97.59%,较基线分别提升37.35%和18.07%。结果证明该架构能高效融合图像与点云,推动跨模态全局定位发展。项目页面与代码将公开。

原文摘要 · Abstract (English)

Image to point cloud global localization is crucial for robot navigation in GNSS-denied environments and has become increasingly important for multi-robot map fusion and urban asset management. The modality gap between images and point clouds poses significant challenges for cross-modality fusion. Current cross-modality global localization solutions either require modality unification, which leads to information loss, or rely on engineered training schemes to encode multi-modality features, which often lack feature alignment and relation consistency. To address these limitations, we propose, SaliencyI2PLoc, a novel contrastive learning based architecture that fuses the saliency map into feature aggregation and maintains the feature relation consistency on multi-manifold spaces. To alleviate the pre-process of data mining, the contrastive learning framework is applied which efficiently achieves cross-modality feature mapping. The context saliency-guided local feature aggregation module is designed, which fully leverages the contribution of the stationary information in the scene generating a more representative global feature. Furthermore, to enhance the cross-modality feature alignment during contrastive learning, the consistency of relative relationships between samples in different manifold spaces is also taken into account. Experiments conducted on urban and highway scenario datasets demonstrate the effectiveness and robustness of our method. Specifically, our method achieves a Recall@1 of 78.92% and a Recall@20 of 97.59% on the urban scenario evaluation dataset, showing an improvement of 37.35% and 18.07%, compared to the baseline method. This demonstrates that our architecture efficiently fuses images and point clouds and represents a significant step forward in cross-modality global localization. The project page and code will be released.

跨模态定位对比学习点云显著性图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。