通过显式建模跨视角对应关系,提升遥感与街景图像定位精度
CLNet: Cross-View Correspondence Makes a Stronger Geo-Localizationer
- 设计三模块框架,显式学习跨视角特征对应关系
- 在四个公开数据集上达到当前最佳定位性能
- 适合需要高精度地理定位与模型可解释性的研究者
基于图像检索的跨视角地理定位(IRCVGL)旨在匹配视角差异显著的图像,如卫星图与街景图。现有方法主要依赖鲁棒的全局表示或隐式特征对齐,难以建模影响定位准确性的显式空间对应关系。本文提出一种新型的对应感知特征优化框架CLNet,显式弥合不同视角间的语义与几何差距。CLNet将视角对齐过程分解为三个可学习且互补的模块:神经对应图(NCM)通过潜在对应场空间对齐跨视角特征;非线性嵌入转换器(NEC)利用MLP变换实现视角间特征重映射;全局特征重校准(GFR)模块则根据学习到的空间线索重新加权关键特征通道。该方法能联合捕捉高层语义与细粒度对齐。在四个公开基准数据集CVUSA、CVACT、VIGOR和University-1652上的大量实验表明,所提CLNet不仅达到当前最优性能,还具备更强的可解释性与泛化能力。
原文摘要 · Abstract (English)
Image retrieval-based cross-view geo-localization (IRCVGL) aims to match images captured from significantly different viewpoints, such as satellite and street-level images. Existing methods predominantly rely on learning robust global representations or implicit feature alignment, which often fail to model explicit spatial correspondences crucial for accurate localization. In this work, we propose a novel correspondence-aware feature refinement framework, termed CLNet, that explicitly bridges the semantic and geometric gaps between different views. CLNet decomposes the view alignment process into three learnable and complementary modules: a Neural Correspondence Map (NCM) that spatially aligns cross-view features via latent correspondence fields; a Nonlinear Embedding Converter (NEC) that remaps features across perspectives using an MLP-based transformation; and a Global Feature Recalibration (GFR) module that reweights informative feature channels guided by learned spatial cues. The proposed CLNet can jointly capture both high-level semantics and fine-grained alignments. Extensive experiments on four public benchmarks, CVUSA, CVACT, VIGOR, and University-1652, demonstrate that our proposed CLNet achieves state-of-the-art performance while offering better interpretability and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。