通过动态路由实现跨视角图像定位,提升匹配精度与效率
Learnable Query Aggregation with KV Routing for Cross-view Geo-localisation
- 引入可学习的键值路由机制,动态选择专家子空间处理异构特征
- 在University-1652和SUES-200上以更少参数达到领先性能
- 适合关注跨视角地理定位与轻量化模型设计的研究者
跨视角地理定位(CVGL)旨在通过匹配大规模数据库中的图像来估计查询图像的地理位置。然而,视角差异显著增加了特征聚合与对齐的难度。为此,我们提出一种新型CVGL系统,包含三项关键改进:首先,采用DINOv2骨干网络并结合卷积适配器微调,增强模型对跨视角变化的适应性;其次,设计多尺度通道重分配模块,强化空间表示的多样性和稳定性;最后,提出改进的聚合模块,将多专家(MoE)路由引入交叉注意力框架中,动态为键和值选择专家子空间,实现对异构输入域的自适应处理。在University-1652和SUES-200数据集上的大量实验表明,该方法以更少训练参数取得具有竞争力的性能。
原文摘要 · Abstract (English)
Cross-view geo-localisation (CVGL) aims to estimate the geographic location of a query image by matching it with images from a large-scale database. However, the significant view-point discrepancies present considerable challenges for effective feature aggregation and alignment. To address these challenges, we propose a novel CVGL system that incorporates three key improvements. Firstly, we leverage the DINOv2 backbone with a convolution adapter fine-tuning to enhance model adaptability to cross-view variations. Secondly, we propose a multi-scale channel reallocation module to strengthen the diversity and stability of spatial representations. Finally, we propose an improved aggregation module that integrates a Mixture-of-Experts (MoE) routing into the feature aggregation process. Specifically, the module dynamically selects expert subspaces for the keys and values in a cross-attention framework, enabling adaptive processing of heterogeneous input domains. Extensive experiments on the University-1652 and SUES-200 datasets demonstrate that our method achieves competitive performance with fewer trained parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。