arXiv:2501.07194cs.CV2025-01中稿 · ICASSP 2025被引 14

针对跨视角定位中视点差异问题,提出视点感知注意力模型提升定位精度。

VAGeo: View-specific Attention for Cross-View Object Geo-Localization

  • 设计视点专用位置编码,区分地面与无人机视角特征。
  • 引入通道-空间混合注意力机制,增强特征判别力。
  • 在标准数据集上显著提升定位准确率,尤其对无人机视角效果明显。

跨视角物体地理定位(CVOGL)旨在将地面或无人机拍摄的查询图像中的目标,在卫星图像中准确定位。现有方法通常等同处理地面与无人机视角图像,忽视了其固有的视点差异以及查询图像与卫星参考图像之间的空间相关性。为此,本文提出一种新型视点专用注意力地理定位方法(VAGeo),包含两个关键模块:视点专用位置编码(VSPE)模块和通道-空间混合注意力(CSHA)模块。在对象层面,根据地面与无人机视角的不同特性,设计视点专用的位置编码,以更准确地识别查询图像中的点击目标。在特征层面,通过同时结合通道注意力与空间注意力机制,实现混合注意力学习,以提取更具判别性的特征。大量实验结果表明,所提出的VAGeo在CVOGL数据集上取得显著性能提升:地面视角的[email protected]/[email protected]从45.43%/42.24%提升至48.21%/45.22%;无人机视角则从61.97%/57.66%提升至66.19%/61.87%。

原文摘要 · Abstract (English)

Cross-view object geo-localization (CVOGL) aims to locate an object of interest in a captured ground- or drone-view image within the satellite image. However, existing works treat ground-view and drone-view query images equivalently, overlooking their inherent viewpoint discrepancies and the spatial correlation between the query image and the satellite-view reference image. To this end, this paper proposes a novel View-specific Attention Geo-localization method (VAGeo) for accurate CVOGL. Specifically, VAGeo contains two key modules: view-specific positional encoding (VSPE) module and channel-spatial hybrid attention (CSHA) module. In object-level, according to the characteristics of different viewpoints of ground and drone query images, viewpoint-specific positional codings are designed to more accurately identify the click-point object of the query image in the VSPE module. In feature-level, a hybrid attention in the CSHA module is introduced by combining channel attention and spatial attention mechanisms simultaneously for learning discriminative features. Extensive experimental results demonstrate that the proposed VAGeo gains a significant performance improvement, i.e., improving [email protected]/[email protected] on the CVOGL dataset from 45.43%/42.24% to 48.21%/45.22% for ground-view, and from 61.97%/57.66% to 66.19%/61.87% for drone-view.

地理定位跨视角注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。