arXiv:2603.02726cs.CV2026-03

提出跨视角地理定位新模型,提升不同角度图像的匹配精度。

Cross-view geo-localization, Image retrieval, Multiscale geometric modeling, Frequency domain enhancement

  • 设计空间与频域并行网络,融合全局语义、局部几何和频率稳定性特征
  • 在多个数据集上超越现有方法,尤其在极端视角变化下表现更优
  • 轻量高效,适合部署于资源受限的视觉定位系统

跨视角地理定位(CVGL)旨在建立从显著不同视角拍摄图像之间的空间对应关系,是无卫星信号环境下视觉定位的基础技术。然而,由于严重的几何不对称性、成像域间的纹理不一致以及判别性局部信息的逐步退化,该任务仍具挑战。现有方法多依赖空间域特征对齐,对大尺度视角变化和局部干扰敏感。本文提出空间与频域增强网络(SFDE),利用空间与频域的互补表征。SFDE采用三分支并行架构,分别建模全局语义上下文、局部几何结构及频域统计稳定性,从场景拓扑、多尺度结构模式和频率不变性角度刻画跨域一致性。通过渐进增强与耦合约束,在统一嵌入空间中联合优化互补特征,学习多粒度一致的跨视角表示。大量实验表明,SFDE性能具有竞争力,多数情况下优于当前最优方法,同时保持轻量化与计算高效性。

原文摘要 · Abstract (English)

Cross-view geo-localization (CVGL) aims to establish spatial correspondences between images captured from significantly different viewpoints and constitutes a fundamental technique for visual localization in GNSS-denied environments. Nevertheless, CVGL remains challenging due to severe geometric asymmetry, texture inconsistency across imaging domains, and the progressive degradation of discriminative local information. Existing methods predominantly rely on spatial domain feature alignment, which is inherently sensitive to large scale viewpoint variations and local disturbances. To alleviate these limitations, this paper proposes the Spatial and Frequency Domain Enhancement Network (SFDE), which leverages complementary representations from spatial and frequency domains. SFDE adopts a three branch parallel architecture to model global semantic context, local geometric structure, and statistical stability in the frequency domain, respectively, thereby characterizing consistency across domains from the perspectives of scene topology, multiscale structural patterns, and frequency invariance. The resulting complementary features are jointly optimized in a unified embedding space via progressive enhancement and coupled constraints, enabling the learning of cross-view representations with consistency across multiple granularities. Comprehensive experiments show that SFDE achieves competitive performance and in many cases even surpasses state-of-the-art methods, while maintaining a lightweight and computationally efficient design. {Our code is available at https://github.com/Mashuaishuai669/SFDE

地理定位图像检索频域增强多尺度建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。