统一框架提升跨视角地理定位精度,1米级召回率最高升至39.64%。
A Unified Hierarchical Framework for Fine-grained Cross-view Geo-localization over Large-scale Scenarios
- 将检索与度量定位整合为统一网络,共享参数实现协同学习。
- 在VIGOR数据集上,1米级定位召回率提升超25倍,跨区域达25.58%。
- 适合大规模地理定位场景,尤其关注高精度定位的研究者。
跨视角地理定位是解决大规模定位问题的有力方案,需依次完成检索与度量定位任务以实现细粒度预测。然而,现有方法多为独立设计两个任务的模型,导致协作效率低且训练开销大。本文提出UnifyGeo,一种将检索与度量定位集成于单一网络的统一分层框架。首先采用共享参数的联合学习策略,共同学习多粒度表征,促进两任务相互增强;随后设计由专用损失函数引导的重排序机制,通过提升检索准确率和度量定位参考质量来增强整体性能。大量实验表明,UnifyGeo在任务孤立与关联设置下均显著优于现有方法。尤其在挑战性VIGOR基准上,1米级定位召回率从1.53%提升至39.64%(同区),从0.43%提升至25.58%(跨区)。代码将公开。
原文摘要 · Abstract (English)
Cross-view geo-localization is a promising solution for large-scale localization problems, requiring the sequential execution of retrieval and metric localization tasks to achieve fine-grained predictions. However, existing methods typically focus on designing standalone models for these two tasks, resulting in inefficient collaboration and increased training overhead. In this paper, we propose UnifyGeo, a novel unified hierarchical geo-localization framework that integrates retrieval and metric localization tasks into a single network. Specifically, we first employ a unified learning strategy with shared parameters to jointly learn multi-granularity representation, facilitating mutual reinforcement between these two tasks. Subsequently, we design a re-ranking mechanism guided by a dedicated loss function, which enhances geo-localization performance by improving both retrieval accuracy and metric localization references. Extensive experiments demonstrate that UnifyGeo significantly outperforms the state-of-the-arts in both task-isolated and task-associated settings. Remarkably, on the challenging VIGOR benchmark, which supports fine-grained localization evaluation, the 1-meter-level localization recall rate improves from 1.53\% to 39.64\% and from 0.43\% to 25.58\% under same-area and cross-area evaluations, respectively. Code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。