arXiv:2606.28371cs.CV2026-06

用语义森林提升地表点云与卫星图像跨视图定位精度

GeoISF: Instance Semantic Forest Inspired Large-Scale Cross-View Geo-Localization via Ground LiDAR-to-Satellite Image

  • 构建基于词典的实例语义森林,融合多帧语义树增强表征
  • KITTI数据集上定位准确率比现有方法高13.22倍(R@10)
  • 适合大规模地理定位任务,尤其在模态差异大的场景

在给定查询地表视角点云的情况下,实现大尺度卫星图像中的定位仍具挑战性。现有激光雷达到图像的跨视图定位方法因语义对齐有限及点云与卫星图像之间的模态差距,在大尺度场景中表现不佳。本文提出名为GeoISF的大规模激光雷达到图像地理定位流程。GeoISF利用WordNet构建实例语义森林,通过整合多帧的语义树增强时间语义表示与区分能力。借助环境语义表示作为共享媒介,有效弥合模态差距并提升语义匹配精度。大量实验表明,GeoISF在大规模跨视图定位中表现卓越,在KITTI数据集上的R@10指标相比平行方法提升13.22倍。该方法解决了大尺度激光雷达到图像跨视图定位的现存短板,为计算与精度挑战提供了稳健方案。代码将开源供学术界使用。

原文摘要 · Abstract (English)

The problem of localization on a large-scale satellite image given a frame of query ground view point clouds remains challenging. Existing LiDAR-to-image cross-view localization methods struggle in large-scale scenarios due to limited semantic alignment and the modality gap between point clouds and satellite images. This paper introduces the large-scale LiDAR-to-image geo-localization pipeline called GeoISF. GeoISF introduces an instance semantic forest constructed using WordNet, which enhances temporal semantic representation and discriminative power by integrating semantic trees from multiple frames. By leveraging environmental semantic representation as a shared medium, GeoISF effectively bridges the modality gap and improves semantic matching accuracy. Extensive experiments demonstrate the superior performance of GeoISF in large-scale cross-view localization, 13.22 times better than the parallel LiDAR-to-image method in the R@10 metric on the KITTI dataset. The proposed method addresses the existing gap in large-scale LiDAR-to-image cross-view localization, offering a robust solution to the computational and accuracy challenges inherent in such scenarios. We will release the code as an open-source resource available online for the broader research community.

地理定位跨模态语义建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。