用距离感知排名提升全球图像定位准确率
GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization
- 基于视觉语言模型联合编码查询与候选图像,建模空间关系
- 引入多阶距离损失,同时优化绝对与相对位置排序
- 首个专为地理排序设计的数据集,适合图像定位研究者
全球图像地理定位——从任意地点拍摄的图像中预测其GPS坐标——因地区间视觉内容差异巨大而面临根本挑战。现有方法采用检索候选后选择最优匹配的两阶段流程,但通常依赖简单相似性启发式和点级监督,无法建模候选间的空间关系。本文提出GeoRanker,一种距离感知排名框架,利用大规模视觉语言模型联合编码查询-候选交互并预测地理接近度。此外,我们引入多阶距离损失,对绝对与相对距离进行排序,使模型能够推理结构化空间关系。为支持该任务,我们构建了首个专为地理排序设计的数据集GeoRanking,包含多模态候选信息。GeoRanker在两个主流基准IM2GPS3K和YFCC4K上达到领先性能,显著优于当前最佳方法。
原文摘要 · Abstract (English)
Worldwide image geolocalization-the task of predicting GPS coordinates from images taken anywhere on Earth-poses a fundamental challenge due to the vast diversity in visual content across regions. While recent approaches adopt a two-stage pipeline of retrieving candidates and selecting the best match, they typically rely on simplistic similarity heuristics and point-wise supervision, failing to model spatial relationships among candidates. In this paper, we propose GeoRanker, a distance-aware ranking framework that leverages large vision-language models to jointly encode query-candidate interactions and predict geographic proximity. In addition, we introduce a multi-order distance loss that ranks both absolute and relative distances, enabling the model to reason over structured spatial relationships. To support this, we curate GeoRanking, the first dataset explicitly designed for geographic ranking tasks with multimodal candidate information. GeoRanker achieves state-of-the-art results on two well-established benchmarks (IM2GPS3K and YFCC4K), significantly outperforming current best methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。