动态路由选择检索或生成,提升全球图像定位精度。
GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization
- 根据图像内容自动选择最优的检索或生成定位方式。
- 在IM2GPS3k和YFCC4k上显著优于现有方法。
- 首个专用于训练路由策略的大规模数据集。
全球图像地理定位旨在为地球任意位置拍摄的图像预测精确的GPS坐标,由于视觉与地理多样性大,任务极具挑战。现有方法主要分为两类:基于检索的方法通过匹配查询与参考数据库实现定位;基于生成的方法则利用大视觉语言模型(LVLM)直接预测坐标。我们发现两者误差模式互补:检索擅长细粒度实例匹配,生成具备更强语义推理能力。这种异质性表明单一范式无法普适。为此,我们提出GeoRouter,一种动态路由框架,可自适应地将每张查询图像分配至最优范式。该框架以LVLM为骨干,分析视觉内容并作出路由决策。为优化路由策略,我们设计距离感知偏好目标,将两范式间的距离差距转化为连续监督信号,显式反映相对性能差异。此外,我们构建了首个多范式独立预测的大型数据集GeoRouting,用于训练路由策略。在IM2GPS3k和YFCC4k上的大量实验表明,GeoRouter显著超越现有最先进方法。
原文摘要 · Abstract (English)
Worldwide image geolocalization aims to predict precise GPS coordinates for images captured anywhere on Earth, which is challenging due to the large visual and geographic diversity. Recent methods mainly follow two paradigms: retrieval-based approaches that match queries against a reference database, and generation-based approaches that directly predict coordinates using Large Vision-Language Models (LVLMs). However, we observe distinct error profiles between them: retrieval excels at fine-grained instance matching, while generation offers robust semantic reasoning. This complementary heterogeneity suggests that no single paradigm is universally superior. To harness this potential, we propose GeoRouter, a dynamic routing framework that adaptively assigns each query to the optimal paradigm. GeoRouter leverages an LVLM backbone to analyze visual content and provide routing decisions. To optimize GeoRouter, we introduce a distance-aware preference objective that converts the distance gap between paradigms into a continuous supervision signal, explicitly reflecting relative performance differences. Furthermore, we construct GeoRouting, the first large-scale dataset tailored for training routing policies with independent paradigm predictions. Extensive experiments on IM2GPS3k and YFCC4k demonstrate that GeoRouter significantly outperforms state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。