arXiv:2604.04357cs.CV2026-04

用地理距离加权改进图像文本对比学习,提升街景定位精度

Spatially-Weighted CLIP for Street-View Geo-localization

  • 引入地理位置文本表示和距离感知软标签,建模空间相关性
  • 在多城市数据集上定位准确率显著提升,长尾误差减少
  • 适合需要精准地理定位的视觉语言模型研究者

本文提出空间加权CLIP(SW-CLIP),一种将空间自相关性显式融入视觉语言对比学习的街景地理定位新框架。与传统CLIP方法将所有非匹配样本视为同等负样本不同,SW-CLIP基于托伯勒地理学第一定律,通过测地距离构建距离感知的软监督信号。具体地,采用位置作为文本表示编码地理坐标,并以空间加权软标签替代原有的单热InfoNCE目标;同时引入邻域一致性正则化,保持嵌入空间中的局部空间结构。在多城市数据集上的实验表明,相较于标准CLIP,SW-CLIP显著提升了定位准确性,降低了长尾误差,并增强了空间一致性。结果凸显了从语义对齐转向地理对齐对于鲁棒地理定位的重要性,并为多模态表征学习中融合空间原理提供了通用范式。

原文摘要 · Abstract (English)

This paper proposes Spatially-Weighted CLIP (SW-CLIP), a novel framework for street-view geo-localization that explicitly incorporates spatial autocorrelation into vision-language contrastive learning. Unlike conventional CLIP-based methods that treat all non-matching samples as equally negative, SW-CLIP leverages Tobler's First Law of Geography to model geographic relationships through distance-aware soft supervision. Specifically, we introduce a location-as-text representation to encode geographic positions and replace one-hot InfoNCE targets with spatially weighted soft labels derived from geodesic distance. Additionally, a neighborhood-consistency regularization is employed to preserve local spatial structure in the embedding space. Experiments on a multi-city dataset demonstrate that SW-CLIP significantly improves geo-localization accuracy, reduces long-tail errors, and enhances spatial coherence compared to standard CLIP. The results highlight the importance of shifting from semantic alignment to geographic alignment for robust geo-localization and provide a general paradigm for integrating spatial principles into multimodal representation learning.

地理定位对比学习空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。