利用3D结构先验提升跨视角地理定位的泛化能力
GeoLink: A 3D-aware Framework to Improve Generalization for Cross-view Geo-localization
- 基于无人机图像重建3D点云,为2D特征提供几何约束
- 在多个基准上超越现有方法,跨域和多天气场景表现优异
- 无需改动2D推理流程,适合实际部署的地理定位系统
可泛化的跨视角地理定位旨在无GPS监督下匹配不同视角的同一位置,其核心挑战在于视点变化导致的语义不一致及域偏移下的泛化能力差。现有方法主要依赖2D对应关系,易受跨视角冗余信息干扰,导致表征迁移性不足。为此,我们提出GeoLink,一种面向可泛化跨视角地理定位的3D感知语义一致框架。具体地,我们使用VGGT离线从多视角无人机图像重建场景点云,提供稳定的结构先验。基于这些3D锚点,通过两个互补方式改进2D表征学习:几何感知语义精修模块在3D引导下缓解2D特征中潜在的冗余与视角偏差依赖;统一视角关系蒸馏模块将3D结构关系迁移到2D特征中,增强跨视角对齐,同时保持纯2D推理流程。大量实验表明,GeoLink在多个基准上持续优于当前最优方法,在未见域和多样天气环境下均展现出卓越泛化性能。
原文摘要 · Abstract (English)
Generalizable cross-view geo-localization aims to match the same location across views in unseen regions and conditions without GPS supervision. Its core difficulty lies in severe semantic inconsistency caused by viewpoint variation and poor generalization under domain shift. Existing methods mainly rely on 2D correspondence, but they are easily distracted by redundant shared information across views, leading to less transferable representations. To address this, we propose GeoLink, a 3D-aware semantic-consistent framework for Generalizable cross-view geo-localization. Specifically, we offline reconstruct scene point clouds from multi-view drone images using VGGT, providing stable structural priors. Based on these 3D anchors, we improve 2D representation learning in two complementary ways. A Geometric-aware Semantic Refinement module mitigates potentially redundant and view-biased dependencies in 2D features under 3D guidance. In addition, a Unified View Relation Distillation module transfers 3D structural relations to 2D features, improving cross-view alignment while preserving a 2D-only inference pipeline. Extensive experiments on multiple benchmarks show that GeoLink consistently outperforms state-of-the-art methods and achieves superior generalization across unseen domains and diverse weather environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。