arXiv:2410.04939cs.CV2024-10中稿 · IEEE TITS 2024被引 8

融合图像与点云,提升复杂环境下的场景识别精度与鲁棒性

PRFusion: Toward Effective and Robust Multi-Modal Place Recognition with Image and Point Cloud Fusion

  • 通过流形度量注意力实现图像与点云全局特征融合,无需相机激光雷达标定
  • 在Boreas数据集上相比现有模型提升3.0% AR@1,显著增强环境适应能力
  • 适合自动驾驶、机器人定位等需要高可靠场景识别的工程应用

场景识别在机器人和计算机视觉中至关重要,广泛应用于自动驾驶、建图与定位。其核心是利用查询传感器数据与已知数据库匹配地点。主要挑战在于模型需在环境变化下仍保持高精度。本文提出两种多模态场景识别模型:PRFusion与PRFusion++。PRFusion采用流形度量注意力进行全局特征融合,可在无需相机-LiDAR外参标定条件下有效交互特征。PRFusion++假设外参已知,利用像素-点对应关系在局部窗口增强特征学习。两者均引入神经扩散层,显著提升在恶劣环境中的可靠性。我们在三个大规模基准上验证了模型性能,尤其在高难度Boreas数据集上,相较现有方法取得+3.0 AR@1的显著提升。消融实验进一步验证了方法有效性。代码已开源。

原文摘要 · Abstract (English)

Place recognition plays a crucial role in the fields of robotics and computer vision, finding applications in areas such as autonomous driving, mapping, and localization. Place recognition identifies a place using query sensor data and a known database. One of the main challenges is to develop a model that can deliver accurate results while being robust to environmental variations. We propose two multi-modal place recognition models, namely PRFusion and PRFusion++. PRFusion utilizes global fusion with manifold metric attention, enabling effective interaction between features without requiring camera-LiDAR extrinsic calibrations. In contrast, PRFusion++ assumes the availability of extrinsic calibrations and leverages pixel-point correspondences to enhance feature learning on local windows. Additionally, both models incorporate neural diffusion layers, which enable reliable operation even in challenging environments. We verify the state-of-the-art performance of both models on three large-scale benchmarks. Notably, they outperform existing models by a substantial margin of +3.0 AR@1 on the demanding Boreas dataset. Furthermore, we conduct ablation studies to validate the effectiveness of our proposed methods. The codes are available at: https://github.com/sijieaaa/PRFusion

场景识别多模态融合点云自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。