用文本锚点连接多视角图像与文字,实现更精准的地理定位
GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
- 通过文本描述桥接无人机、街景与卫星图像特征
- 在5万+跨视图数据上训练,定位准确率显著提升
- 适合需要多源地理信息融合的定位研究者
跨视图地理定位通过检索与查询图像视觉匹配的带地理标签参考图像来推断位置。然而,传统以卫星为中心的方法在缺乏高分辨率或最新卫星影像时鲁棒性差,且未能充分利用不同视角(如无人机、卫星、街景)和模态(如语言与图像)间的互补信息。为此,我们提出GeoBridge,一种支持双向跨视图匹配及语言到图像检索的新模型。超越传统卫星主导范式,GeoBridge基于新颖的语义锚机制,通过文本描述连接多视角特征,实现更鲁棒、灵活的定位。为支持该任务,我们构建了首个大规模、跨模态、多视角对齐数据集GeoLoc,包含来自36个国家的超过5万对无人机、街景全景与卫星图像及其文本描述,确保地理与语义对齐。我们在多项任务上进行了广泛评估,实验表明,使用GeoLoc预训练显著提升了GeoBridge的地理定位准确率,并增强了跨域泛化能力与跨模态知识迁移能力。代码、数据集与预训练模型将发布于https://github.com/MiliLab/GeoBridge。
原文摘要 · Abstract (English)
Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date satellite imagery is unavailable. It further underexploits complementary cues across views (\eg, drone, satellite, and street) and modalities (\eg, language and image). To address these challenges, we propose GeoBridge, a novel model that performs bidirectional matching across views and supports language-to-image retrieval. Going beyond traditional satellite-centric formulations, GeoBridge builds on a novel semantic-anchor mechanism that bridges multi-view features through textual descriptions for robust, flexible localization. In support of this task, we construct GeoLoc, the first large-scale, cross-modal, and multi-view aligned dataset comprising over 50,000 pairs of drone, street-view panorama, and satellite images as well as their textual descriptions, collected from 36 countries, ensuring both geographic and semantic alignment. We performed broad evaluations across multiple tasks. Experiments confirm that GeoLoc pre-training markedly improves geo-location accuracy for GeoBridge while promoting cross-domain generalization and cross-modal knowledge transfer. Code, dataset, and pretrained models will be released at https://github.com/MiliLab/GeoBridge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。