构建百万级跨模态地理定位数据集,提升全球范围定位精度。
Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework
- 用大视觉语言模型生成高质量场景描述,增强文本表征。
- 提出物理一致性对比学习框架,显著提升跨模态定位性能。
- 适用于全球导航、应急响应等需高鲁棒性的定位任务。
跨模态地理定位(CMGL)旨在将地面文本描述与带有地理标签的航拍图像匹配,对行人导航和应急响应至关重要。然而现有研究受限于地理覆盖范围窄和场景多样性不足,难以反映全球建筑风格与地形特征的复杂差异。为填补这一空白并推动通用定位技术发展,我们提出CORE,首个百万级全球跨模态地理定位数据集。CORE包含来自六大洲225个地理区域的1,034,786张跨视角图像,涵盖多样环境条件与城市布局。我们利用大视觉语言模型(LVLMs)的零样本推理能力,生成富含判别性线索的高质量场景描述。此外,我们提出物理法则感知网络(PLANET),引入新型对比学习范式,引导文本表示捕捉卫星图像中的内在物理特征。在多区域实验中,PLANET显著优于现有先进方法,确立了鲁棒、全局尺度地理定位的新基准。数据集与源代码将公开于https://github.com/YtH0823/CORE。
原文摘要 · Abstract (English)
Cross-modal Geo-localization (CMGL) matches ground-level text descriptions with geo-tagged aerial imagery, which is crucial for pedestrian navigation and emergency response. However, existing studies are constrained by narrow geographic coverage and simplistic scene diversity, failing to reflect the immense spatial heterogeneity of global architectural styles and topographic features. To bridge this gap and facilitate universal positioning, we introduce CORE, the first million-scale dataset dedicated to global CMGL. CORE comprises 1,034,786 cross-view images sampled from 225 distinct geographic regions across six continents, offering an unprecedented variety of perspectives in varying environmental conditions and urban layouts. We leverage the zero-shot reasoning of Large Vision-Language Models (LVLMs) to synthesize high-quality scene descriptions rich in discriminative cues. Furthermore, we propose a physical-law-aware network (PLANET) for cross-modal geo-localization. PLANET introduces a novel contrastive learning paradigm to guide textual representations in capturing the intrinsic physical signatures of satellite imagery. Extensive experiments across varied geographic regions demonstrate that PLANET significantly outperforms state-of-the-art methods, establishing a new benchmark for robust, global-scale geo-localization. The dataset and source code will be released at https://github.com/YtH0823/CORE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。