arXiv:2606.00784cs.CV2026-06

用语义门控融合与Mamba序列聚合,提升无人机无卫星定位精度

DINO-GFSA: Geo-Localization via Semantic Gated Fusion and Mamba-based Sequential Aggregation

论文配图:DINO-GFSA: Geo-Localization via Semantic Gated Fusion and Mamba-based Sequential Aggregation
图 1 · 摘自论文原文
  • 通过语义门控机制动态融合高层语义与低层空间特征
  • 在DenseUAV数据集上Recall@1达92.38%,比之前最佳提升3.48%
  • 适合需要高精度定位的无人机自主导航场景

跨视图地理定位(CVGL)对无卫星信号环境下的无人机自主定位和目标定位至关重要。然而,在保持细粒度空间细节的同时获取鲁棒语义仍具挑战。为此,我们提出DINO-GFSA框架,采用LoRA适配的DINOv3(ViTL)主干网络实现参数高效、高容量表征。关键创新在于引入语义门控残差融合模块,利用高层语义选择性校准并融合底层空间线索,有效弥合语义鸿沟。此外,设计基于Mamba的序列聚合头,以线性复杂度捕捉长程空间依赖。实验表明,该方法在University-1652和DenseUAV基准上均达到领先性能,尤其在DenseUAV上相较此前最优结果提升3.48% Recall@1。结果验证了DINO-GFSA作为通用、鲁棒的无人机CVGL解决方案的有效性。

原文摘要 · Abstract (English)

Cross-view geo-localization (CVGL) is critical for Unmanned Aerial Vehicle (UAV) self-positioning and target localization in GNSS-denied environments. However, acquiring robust semantics while preserving finegrained spatial details remains challenging. To address this, we propose DINO-GFSA, a framework leveraging a LoRA (Low-Rank Adaptation) adapted DINOv3 (ViTL) backbone for parameter-efficient, high-capacity representation. Crucially, we introduce a Semantic Gated Residual Fusion module, which utilizes high-level semantics to selectively calibrate and integrate low-level spatial cues, effectively bridging the semantic gap. Furthermore, a Mamba-based Sequential Aggregation Head is designed to capture long-range spatial dependencies with linear complexity. Experiments demonstrate state-of-the-art performance on University-1652 and DenseUAV benchmarks, notably surpassing the previous best on DenseUAV by 3.48% on Recall@1. These results validate DINO-GFSA as a generalized, robust solution for UAV CVGL.

地理定位无人机视觉融合Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。