arXiv:2505.18700cs.CVcs.AI2025-05NeurIPS被引 17

用增强推理链提升视觉语言模型的地理定位能力

GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains

  • 通过多阶段推理逐步分析场景、细节和语义特征
  • 在30K图像数据集上实现跨粗粒度与细粒度定位的显著提升
  • 适合需要可解释地理推理的智能导航与地图应用

视觉语言模型在视觉推理任务中表现优异,但地理定位需从图像中提取多粒度视觉线索,并融合外部世界知识进行系统推理。现有方法缺乏稳健的推理机制和可解释性,限制了效果。为此,我们提出地理推理增强(GRE)套件,从数据集、模型和评测三方面构建框架。首先引入高质量的地理定位推理数据集GRE30K,支持细粒度视觉与上下文分析。其次提出GRE模型,采用多阶段推理策略,逐步推断场景属性、局部细节和语义特征,从而更精确缩小地理范围。最后构建地理推理评估基准GREval-Bench,涵盖城市、自然和地标场景,评估粗粒度(如国家、大洲)与细粒度(如城市、街道)定位性能。实验表明,GRE在所有粒度级别均显著优于现有方法,验证了增强推理的VLM在复杂地理推断中的有效性。代码与数据将开源。

原文摘要 · Abstract (English)

Recent advances in Visual Language Models (VLMs) have demonstrated exceptional performance in visual reasoning tasks. However, geo-localization presents unique challenges, requiring the extraction of multigranular visual cues from images and their integration with external world knowledge for systematic reasoning. Current approaches to geo-localization tasks often lack robust reasoning mechanisms and explainability, limiting their effectiveness. To address these limitations, we propose the Geo Reason Enhancement (GRE) Suite, a novel framework that augments VLMs with structured reasoning chains for accurate and interpretable location inference. The GRE Suite is systematically developed across three key dimensions: dataset, model, and benchmark. First, we introduce GRE30K, a high-quality geo-localization reasoning dataset designed to facilitate fine-grained visual and contextual analysis. Next, we present the GRE model, which employs a multi-stage reasoning strategy to progressively infer scene attributes, local details, and semantic features, thereby narrowing down potential geographic regions with enhanced precision. Finally, we construct the Geo Reason Evaluation Benchmark (GREval-Bench), a comprehensive evaluation framework that assesses VLMs across diverse urban, natural, and landmark scenes to measure both coarse-grained (e.g., country, continent) and fine-grained (e.g., city, street) localization performance. Experimental results demonstrate that GRE significantly outperforms existing methods across all granularities of geo-localization tasks, underscoring the efficacy of reasoning-augmented VLMs in complex geographic inference. Code and data will be released at https://github.com/Thorin215/GRE.

地理定位视觉语言模型推理链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。