arXiv:2512.17492cs.CV2025-12中稿 · CVPR

构建跨视角地理地标多模态数据集,推动地理空间理解研究

MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

  • 构建四模态统一标注的地标数据集,支持跨视角检索与定位
  • 包含18.557个地标、197k航拍图、329k地面视图,实现模态间一一对应
  • 验证现有模型在多模态地理任务中表现不足,适合多模态地理研究者

地理空间分析受益于多模态方法,每个地理地点可从多种方式描述(不同视角的图像、文本描述、地理坐标等)。现有基准在模态覆盖上有限,导致模型局限于单一领域,无法充分利用其他地理模态。本文提出多模态地标数据集(MMLandmarks),包含18.557个美国地标,涵盖197,000张高分辨率航拍图像、329,000张地面视角图像、文本信息及地理坐标,所有模态在地标级别一一对应。该数据集支持跨视角地表到卫星检索、地表与卫星地理定位、文本到图像、文本到GPS检索等多种任务。实验表明,当前专用或现成的基础模型难以直接应用于此类多模态地理任务,暴露出多模态数据对更广泛地理理解的必要性。采用类CLIP的简单基线模型,在MMLandmarks上训练后展现出良好的泛化能力。

原文摘要 · Abstract (English)

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current benchmarks have limited coverage across modalities, leading to specialized models that perform well in their respective domains, but do not fully take advantage of other geo-spatial modalities. We introduce the Multi-Modal Landmark dataset (MMLandmarks), a benchmark composed of four modalities: 197k high-resolution aerial images, 329k ground-view images, textual information, and geographic coordinates for 18.557 distinct landmarks in the United States. The MMLandmarks dataset has a one-to-one landmark level correspondence across every modality, which enables training and benchmarking models for various geo-spatial tasks, including cross-view Ground-to-Satellite retrieval, ground and satellite geolocalization, Text-to-Image, and Text-to-GPS retrieval. We show that current specialized and off-the-shelf foundation models cannot be trivially used to solve this variety of geo-spatial tasks, illustrating a gap where multimodal datasets lead to broader geo-spatial understanding. We employ a simple CLIP-inspired baseline that reflects versatility and broad generalization when trained with MMLandmarks.

地理空间多模态数据集跨视角检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。