arXiv:2509.07450cs.CVcs.CL2025-09被引 3

提出可解释的跨视角地理定位模型,统一多视图多模态匹配与推理。

GLEAM: Learning to Match and Explain in Cross-View Geo-Localization

  • 用卫星图像统一对齐多视角多模态数据,提升训练效率。
  • 在多个基准上达到与专用模型相当的定位精度。
  • 结合大语言模型实现匹配理由生成,适合需要可解释性的场景。

跨视角地理定位(CVGL)旨在识别从同一地理位置不同视角拍摄图像间的对应关系。现有方法通常局限于单一视角或模态,且直接视觉匹配缺乏可解释性:仅判断图像是否匹配,无法说明原因。本文提出GLEAM-C,一个基于卫星图像统一对齐多视图多模态的通用CVGL模型,通过优化实现和新颖的两阶段训练策略,实现与先前专用模型相当的精度,同时显著提升训练效率。为解决可解释性问题,进一步提出GLEAM-X任务,结合跨视图对应预测与由多模态大语言模型(MLLMs)驱动的可解释推理。我们使用商业级MLLMs构建双语基准数据集,并通过严格的人工修订完善测试集,系统评估可解释的跨视图推理能力。GLEAM-C与GLEAM-X共同构成完整流程,实现多模态、多视图对齐与可解释匹配分析,推动地理定位向更可信、可理解的方向发展。代码与数据集将公开于https://github.com/Lucky-Lance/GLEAM。

原文摘要 · Abstract (English)

Cross-View Geo-Localization (CVGL) focuses on identifying correspondences between images captured from distinct perspectives of the same geographical location. However, existing CVGL approaches are typically restricted to a single view or modality, and their direct visual matching strategy lacks interpretability: they only determine whether two images correspond, without explaining the rationale behind the match. In this paper, we present GLEAM-C, a foundational CVGL model that unifies multiple views and modalities by aligning them exclusively with satellite imagery. Our framework improves training efficiency through optimized implementation and achieves accuracy comparable to prior modality-specific CVGL models via a novel two-phase training strategy. To address interpretability, we further propose GLEAM-X, a novel task that combines cross-view correspondence prediction with explainable reasoning enabled by multimodal large language models (MLLMs). We construct a bilingual benchmark using commercial MLLMs to generate training and testing data, and refine the test set through rigorous human revision for systematic evaluation of explainable cross-view reasoning. Together, GLEAM-C and GLEAM-X form a comprehensive CVGL pipeline that integrates multi-modal, multi-view alignment with interpretable correspondence analysis, unifying accurate cross-view matching with explainable reasoning and advancing Geo-Localization by enabling models to better Explain And Match. Code and datasets used in this work will be made publicly accessible at https://github.com/Lucky-Lance/GLEAM.

地理定位可解释性多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。