arXiv:2509.12673cs.CVcs.AI2025-09被引 1

用多尺度频域注意力融合提升跨视角地理定位精度

MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization

  • 通过多频段分支块捕捉不同尺度的结构与边缘特征
  • 频域感知空间注意力模块有效抑制背景噪声干扰
  • 在University-1652等三个数据集上表现优异,适合无人机定位场景

跨视角地理定位旨在通过将查询图像与图库图像匹配来确定其地理位置。该任务因视角变化导致的物体外观差异大,且难以提取判别性特征而具有挑战性。现有方法通常依赖特征图分割提取特征,忽视了空间与语义信息。为此,本文提出基于EVA02的多尺度频域注意力融合(MFAF)方法,包含多频段分支块(MFB)和频域感知空间注意力(FSA)模块。MFB块在多尺度下有效捕获低频结构特征与高频边缘细节,增强不同视角下的特征表示一致性与鲁棒性。FSA模块自适应聚焦于频域特征的关键区域,显著降低背景噪声与视角变化带来的干扰。在University-1652、SUES-200和Dense-UAV等主流基准上的大量实验表明,MFAF在无人机定位与导航任务中均取得竞争力表现。

原文摘要 · Abstract (English)

Cross-view geo-localization aims to determine the geographical location of a query image by matching it against a gallery of images. This task is challenging due to the significant appearance variations of objects observed from variable views, along with the difficulty in extracting discriminative features. Existing approaches often rely on extracting features through feature map segmentation while neglecting spatial and semantic information. To address these issues, we propose the EVA02-based Multi-scale Frequency Attention Fusion (MFAF) method. The MFAF method consists of Multi-Frequency Branch-wise Block (MFB) and the Frequency-aware Spatial Attention (FSA) module. The MFB block effectively captures both low-frequency structural features and high-frequency edge details across multiple scales, improving the consistency and robustness of feature representations across various viewpoints. Meanwhile, the FSA module adaptively focuses on the key regions of frequency features, significantly mitigating the interference caused by background noise and viewpoint variability. Extensive experiments on widely recognized benchmarks, including University-1652, SUES-200, and Dense-UAV, demonstrate that the MFAF method achieves competitive performance in both drone localization and drone navigation tasks.

地理定位频域注意力无人机导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。