一个模型搞定所有视角的地理定位,无需多个模型切换。
SinGeo: Unlock Single Model's Potential for Robust Cross-View Geo-Localization
- 采用双分支判别学习+课程训练,提升模型跨视角识别能力。
- 在四个数据集上达领先水平,极端视角下表现优于专门训练模型。
- 方法简单有效,适合需要稳定定位的导航与遥感应用。
尽管近期进展显著,鲁棒的跨视角地理定位(CVGL)仍具挑战性。现有方法依赖特定视场角(FoV)训练,模型在未见视场角和未知朝向下性能急剧下降,需部署多个模型覆盖不同变化。虽有研究尝试随机化视场角进行动态训练,但未能实现对多样条件的鲁棒性——隐含假设所有视场角难度相同。为此,我们提出SinGeo,一种简洁而强大的框架,使单个模型即可实现鲁棒的跨视角地理定位,无需额外模块或显式变换。SinGeo采用双判别学习架构,增强地面与卫星分支内部的判别能力,并首次引入课程学习策略实现鲁棒CVGL。在四个基准数据集上的广泛评估显示,SinGeo在多种条件下达到当前最优(SOTA)性能,尤其在极端视场角下超越专为该场景训练的方法。此外,SinGeo具备跨架构迁移能力。我们还提出一种一致性评估方法,定量衡量模型在不同视角下的稳定性,为未来鲁棒性研究提供可解释视角。代码将在接受后公开。
原文摘要 · Abstract (English)
Robust cross-view geo-localization (CVGL) remains challenging despite the surge in recent progress. Existing methods still rely on field-of-view (FoV)-specific training paradigms, where models are optimized under a fixed FoV but collapse when tested on unseen FoVs and unknown orientations. This limitation necessitates deploying multiple models to cover diverse variations. Although studies have explored dynamic FoV training by simply randomizing FoVs, they failed to achieve robustness across diverse conditions -- implicitly assuming all FoVs are equally difficult. To address this gap, we present SinGeo, a simple yet powerful framework that enables a single model to realize robust cross-view geo-localization without additional modules or explicit transformations. SinGeo employs a dual discriminative learning architecture that enhances intra-view discriminability within both ground and satellite branches, and is the first to introduce a curriculum learning strategy to achieve robust CVGL. Extensive evaluations on four benchmark datasets reveal that SinGeo sets state-of-the-art (SOTA) results under diverse conditions, and notably outperforms methods specifically trained for extreme FoVs. Beyond superior performance, SinGeo also exhibits cross-architecture transferability. Furthermore, we propose a consistency evaluation method to quantitatively assess model stability under varying views, providing an explainable perspective for understanding and advancing robustness in future CVGL research. Codes will be available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。