arXiv:2411.12431cs.CV2024-11中稿 · IEEE JSTARS被引 43

用大模型+新策略提升全球街景定位精度

CV-Cities: Advancing Cross-View Geo-Localization in Global Cities

  • 融合DINOv2与特征混合器,设计对称InfoNCE损失
  • 在多数据集上超越现有方法,全球定位准确率显著提升
  • 适合做跨视角地理定位研究的学者和工程人员

跨视图地理定位(CVGL)在无GNSS信号场景中至关重要,但受限于视角差异大、场景复杂及全球定位需求。本文提出新框架,结合视觉基础模型DINOv2与先进特征混合器,引入对称InfoNCE损失,并采用近邻采样与动态相似性采样策略,显著提升定位精度。实验表明该框架在多个公开及自建数据集上均优于现有方法。为进一步推动全局性能,我们构建了CV-Cities数据集,包含223,736对地面-卫星图像,覆盖六大洲十六座城市,涵盖多样复杂场景,为CVGL提供挑战性基准。基于该数据集训练的模型在多个测试城市中表现出高精度与强泛化能力。代码与数据已开源。

原文摘要 · Abstract (English)

Cross-view geo-localization (CVGL), which involves matching and retrieving satellite images to determine the geographic location of a ground image, is crucial in GNSS-constrained scenarios. However, this task faces significant challenges due to substantial viewpoint discrepancies, the complexity of localization scenarios, and the need for global localization. To address these issues, we propose a novel CVGL framework that integrates the vision foundational model DINOv2 with an advanced feature mixer. Our framework introduces the symmetric InfoNCE loss and incorporates near-neighbor sampling and dynamic similarity sampling strategies, significantly enhancing localization accuracy. Experimental results show that our framework surpasses existing methods across multiple public and self-built datasets. To further improve globalscale performance, we have developed CV-Cities, a novel dataset for global CVGL. CV-Cities includes 223,736 ground-satellite image pairs with geolocation data, spanning sixteen cities across six continents and covering a wide range of complex scenarios, providing a challenging benchmark for CVGL. The framework trained with CV-Cities demonstrates high localization accuracy in various test cities, highlighting its strong globalization and generalization capabilities. Our datasets and codes are available at https://github.com/GaoShuang98/CVCities.

地理定位跨视图匹配大模型应用多模态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。