arXiv:2506.23077cs.CV2025-06被引 1

提出动态对比学习框架,提升跨视角地理定位的层次化检索精度。

Dynamic Contrastive Learning for Hierarchical Retrieval: A Case Study of Distance-Aware Cross-View Geo-Localization

  • 设计动态对比学习,按层级空间边界逐步对齐特征表示。
  • 在首个带距离标注的校园数据集上,显著提升定位准确率。
  • 适合关注城市级视觉定位与多尺度特征对齐的研究者。

现有基于深度学习的跨视角地理定位方法主要关注跨域图像匹配精度,而非全面捕捉目标周围上下文信息或降低定位误差成本。为系统研究这一距离感知跨视角地理定位(DACVGL)问题,我们构建了首个包含多视角影像与精确距离标注的基准数据集DA-Campus,覆盖三个空间分辨率。基于该数据集,我们将DACVGL建模为跨域的层次化检索问题。研究发现,由于建筑间空间关系的固有复杂性,该问题只能通过对比学习范式解决,而非传统度量学习。为此,我们提出动态对比学习(DyCL)框架,根据层级空间边界逐步对齐特征表示。大量实验表明,DyCL与现有多尺度度量学习方法高度互补,在层次化检索性能和整体跨视角地理定位准确率上均有显著提升。代码与数据集已公开于https://github.com/anocodetest1/DyCL。

原文摘要 · Abstract (English)

Existing deep learning-based cross-view geo-localization methods primarily focus on improving the accuracy of cross-domain image matching, rather than enabling models to comprehensively capture contextual information around the target and minimize the cost of localization errors. To support systematic research into this Distance-Aware Cross-View Geo-Localization (DACVGL) problem, we construct Distance-Aware Campus (DA-Campus), the first benchmark that pairs multi-view imagery with precise distance annotations across three spatial resolutions. Based on DA-Campus, we formulate DACVGL as a hierarchical retrieval problem across different domains. Our study further reveals that, due to the inherent complexity of spatial relationships among buildings, this problem can only be addressed via a contrastive learning paradigm, rather than conventional metric learning. To tackle this challenge, we propose Dynamic Contrastive Learning (DyCL), a novel framework that progressively aligns feature representations according to hierarchical spatial margins. Extensive experiments demonstrate that DyCL is highly complementary to existing multi-scale metric learning methods and yields substantial improvements in both hierarchical retrieval performance and overall cross-view geo-localization accuracy. Our code and benchmark are publicly available at https://github.com/anocodetest1/DyCL.

地理定位对比学习层次检索多视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。