arXiv:2509.21573cs.CVcs.AI2025-09被引 3

通过空间相关性发现难负样本,提升图像地理定位精度

Enhancing Contrastive Learning for Geolocalization by Discovering Hard Negatives on Semivariograms

  • 用半变异图建模视觉相似性与地理距离的关系
  • 在OSV5M数据集上定位准确率显著提升,细粒度效果更优
  • 适合做高精度全球地理定位的科研与工程人员

基于图像的全球尺度地理定位因环境多样、场景视觉相似且缺乏明显地标而极具挑战。尽管对比学习通过对齐街景图像与位置特征表现出良好性能,却忽视了地理空间中的潜在空间依赖关系。这导致其难以处理假负样本(视觉和地理均相近但被标为负例)以及难负样本(视觉相似但地理距离远)。为此,本文提出一种空间正则化的对比学习策略,引入半变异图(semivariogram),将特征空间距离与地理距离关联,捕捉空间相关性下的预期视觉差异。基于拟合的半变异图,定义特定空间距离下的预期视觉不相似度作为参考,用于识别难负样本与假负样本。该方法集成至GeoCLIP,在OSV5M数据集上验证,结果显示显式建模空间先验能有效提升图像地理定位性能,尤其在细粒度定位中表现突出。

原文摘要 · Abstract (English)

Accurate and robust image-based geo-localization at a global scale is challenging due to diverse environments, visually ambiguous scenes, and the lack of distinctive landmarks in many regions. While contrastive learning methods show promising performance by aligning features between street-view images and corresponding locations, they neglect the underlying spatial dependency in the geographic space. As a result, they fail to address the issue of false negatives -- image pairs that are both visually and geographically similar but labeled as negatives, and struggle to effectively distinguish hard negatives, which are visually similar but geographically distant. To address this issue, we propose a novel spatially regularized contrastive learning strategy that integrates a semivariogram, which is a geostatistical tool for modeling how spatial correlation changes with distance. We fit the semivariogram by relating the distance of images in feature space to their geographical distance, capturing the expected visual content in a spatial correlation. With the fitted semivariogram, we define the expected visual dissimilarity at a given spatial distance as reference to identify hard negatives and false negatives. We integrate this strategy into GeoCLIP and evaluate it on the OSV5M dataset, demonstrating that explicitly modeling spatial priors improves image-based geo-localization performance, particularly at finer granularity.

地理定位对比学习半变异图空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。