arXiv:2605.10345cs.CV2026-05

用视觉大模型适配解决航拍与卫星图像的几何差异问题

BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization

论文配图:BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization
图 1 · 摘自论文原文
  • 通过多粒度特征增强和频域结构聚合,小成本提升跨视角特征对齐
  • 在University-1652和SUES-200数据集上达到顶尖定位精度
  • 适合关注跨视角地理定位与高效模型适配的研究者

跨视角图像(如无人机与卫星视图)间的几何差异显著增加了跨视角地理定位(CVGL)的难度,该任务旨在通过图像检索获取图像地理位置。为进一步提升CVGL性能,本文提出一种基于视觉基础模型(如DINOv3)的参数高效适配框架BGG,用于弥合跨视角图像间的几何差距。BGG不仅有效利用视觉基础模型的通用视觉表征,捕捉跨视角图像中的鲁棒一致特征,还借助其泛化能力显著提升定位性能。框架主要包括多粒度特征增强适配器(MFEA)和频域感知结构聚合(FASA)模块。MFEA通过多层级空洞卷积增强特征的尺度适应性与视角鲁棒性,以低训练成本有效弥合跨视角几何差异。同时,针对[CLS]标记缺乏空间细节的问题,FASA模块在频域调制补丁标记,并自适应聚合局部结构特征以增强。最终,BGG将增强的局部特征与[CLS]标记融合,实现更精准的地理定位。在University-1652和SUES-200数据集上的大量实验表明,BGG相比现有方法具有显著优势,以低训练成本实现了当前最优的定位性能。

原文摘要 · Abstract (English)

Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization (CVGL), which aims to acquire the geolocation of images by image retrieval. To further enhance the CVGL performance, this paper proposes a parameter-efficient adaptation framework for bridging the geometric gap across images based on the vision foundation model (VFM) (e.g., DINOv3), termed BGG. BGG not only effectively leverages the general visual representations of VFM and captures the robust and consistent features from cross-view images, but also utilizes the generalization capabilities of the VFM, significantly improving the CVGL performance. It mainly contains a Multi-granularity Feature Enhancement Adapter (MFEA) and a Frequency-Aware Structural Aggregation (FASA) module. Specifically, MFEA enhances the scale adaptability and viewpoint robustness of features by multi-level dilated convolutions, effectively bridging the cross-view geometric gap with small training costs. Additionally, considering the [CLS] token lacks spatial details for precise image retrieval and localization, the FASA module modulates patch tokens in the frequency domain and performs adaptive aggregation for local structural feature enhancement. Finally, BGG fuses the enhanced local features with the [CLS] token for more accurate CVGL. Extensive experiments on University-1652 and SUES-200 datasets demonstrate that BGG has significant advantages over other methods and achieves state-of-the-art localization performance with low training costs.

地理定位跨视角匹配视觉模型特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。