arXiv:2412.14819cs.CV2024-12中稿 · TGRS 2025被引 37

提出轻量级网络MEAN,提升跨视角地理定位精度与效率

Multi-Level Embedding and Alignment Network with Consistency and Invariance Learning for Cross-View Geo-Localization

  • 分层增强+跨域对齐,实现多层级特征融合与一致性学习
  • 参数减少62.17%,计算量降低70.99%,性能仍领先于顶尖模型
  • 适合资源受限场景下的高精度地理定位应用

跨视角地理定位(CVGL)旨在通过检索最相似的带地理标签卫星图像来确定无人机图像的位置。然而,平台间成像差异大、视角变化显著,导致现有方法难以有效关联跨视角特征并提取一致且不变的特征。此外,提升性能常伴随计算与存储开销增加。为此,本文提出轻量级增强对齐网络MEAN,采用渐进式多层次增强策略、全局到局部关联及跨域对齐机制,实现多层级特征通信,有效建立不同层级间的特征连接,学习鲁棒的跨视角一致映射与模态不变特征。同时,结合浅层主干网络与轻量分支设计,显著降低参数量与计算复杂度。在University-1652和SUES-200数据集上的实验表明,相较当前最优模型,MEAN参数量减少62.17%,计算复杂度降低70.99%,同时保持甚至超越其性能表现。代码与模型将公开于https://github.com/ISChenawei/MEAN。

原文摘要 · Abstract (English)

Cross-View Geo-Localization (CVGL) involves determining the localization of drone images by retrieving the most similar GPS-tagged satellite images. However, the imaging gaps between platforms are often significant and the variations in viewpoints are substantial, which limits the ability of existing methods to effectively associate cross-view features and extract consistent and invariant characteristics. Moreover, existing methods often overlook the problem of increased computational and storage requirements when improving model performance. To handle these limitations, we propose a lightweight enhanced alignment network, called the Multi-Level Embedding and Alignment Network (MEAN). The MEAN network uses a progressive multi-level enhancement strategy, global-to-local associations, and cross-domain alignment, enabling feature communication across levels. This allows MEAN to effectively connect features at different levels and learn robust cross-view consistent mappings and modality-invariant features. Moreover, MEAN adopts a shallow backbone network combined with a lightweight branch design, effectively reducing parameter count and computational complexity. Experimental results on the University-1652 and SUES-200 datasets demonstrate that MEAN reduces parameter count by 62.17% and computational complexity by 70.99% compared to state-of-the-art models, while maintaining competitive or even superior performance. Our code and models will be released on https://github.com/ISChenawei/MEAN.

地理定位跨视角匹配轻量化模型特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。