用分层地理嵌入与语义融合提升图像全球定位精度
GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings
- 构建分层地理嵌入,实现多尺度地理空间表示
- 在5个基准上22/25指标达新最佳,显著超越现有方法
- 适合对高精度图像定位、地理信息融合感兴趣的开发者
全球视觉地理定位旨在仅凭图像视觉内容确定其在地球上的位置。尽管近期取得进展,但地理坐标的低维特性使得学习表达性强的地理表征仍具挑战。本文将全局定位问题建模为查询图像的视觉表示与学习到的地理表示之间的对齐。提出显式建模世界为分层地理嵌入结构,实现地理空间的分布式、多尺度表示。同时引入语义融合模块,通过潜在交叉注意力高效融合外观特征与语义分割图,生成更鲁棒的视觉表示。在五个广泛使用的地理定位基准上实验表明,该方法在25项报告指标中的22项达到新最优。消融研究显示,性能提升主要源于所提出的地理表示和语义融合机制。
原文摘要 · Abstract (English)
Worldwide visual geo-localization aims to determine the geographic location of an image anywhere on Earth using only its visual content. Despite recent progress, learning expressive representations of geographic space remains challenging due to the inherently low-dimensional nature of geographic coordinates. We formulate global geo-localization as aligning the visual representation of a query image with a learned geographic representation. Our approach explicitly models the world as a hierarchy of learned geographic embeddings, enabling a distributed and multi-scale representation of geographic space. In addition, we introduce a semantic fusion module that efficiently integrates appearance features with semantic segmentation through latent cross-attention, producing a more robust visual representation for localization. Experiments on five widely used geo-localization benchmarks demonstrate that our method achieves new state-of-the-art results on 22 of 25 reported metrics. Ablation studies show that these improvements are primarily driven by the proposed geographic representation and semantic fusion mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。