arXiv:2506.14674cs.CV2025-06NeurIPS被引 30

用大模型让地图定位更会推理,提升复杂场景下的准确性。

Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models

  • 构建新数据集MP16-Reason,用社交媒体图像增强场景多样性。
  • 提出GLOBE方法,在定位与推理上同时提升,准确率显著优于现有模型。
  • 适合关注可解释定位、多模态推理的研究者和应用开发者。

以往图像地理定位方法多将其视为分类或检索任务,依赖黑箱决策且缺乏可解释性。大视觉语言模型(LVLM)的兴起使定位任务转向基于视觉线索的推理驱动模式。然而,数据层面现有推理数据集以街景为主,场景多样性和视角受限;建模层面当前方法主要依赖监督微调,推理能力提升有限。为此,我们提出新流程:利用多样化社交媒体图像构建面向推理的地理定位数据集MP16-Reason;引入GLOBE(Group-relative policy optimization for Localizability assessment and Optimized visual-cue reasoning),通过任务特异性奖励联合优化可定位性评估、视觉线索推理与定位精度,实现双目标增强。定性与定量结果表明,GLOBE在多种复杂视觉场景中优于现有开源LVLM,且生成更具洞察力和可解释性的推理路径。数据与代码已公开于https://github.com/lingli1996/GLOBE。

原文摘要 · Abstract (English)

Previous methods for image geo-localization have typically treated the task as either classification or retrieval, often relying on black-box decisions that lack interpretability. The rise of large vision-language models (LVLMs) has enabled a rethinking of geo-localization as a reasoning-driven task grounded in visual cues. However, two major challenges persist. On the data side, existing reasoning-focused datasets are primarily based on street-view imagery, offering limited scene diversity and constrained viewpoints. On the modeling side, current approaches predominantly rely on supervised fine-tuning, which yields only marginal improvements in reasoning capabilities. To address these challenges, we propose a novel pipeline that constructs a reasoning-oriented geo-localization dataset, MP16-Reason, using diverse social media images. We introduce GLOBE, Group-relative policy optimization for Localizability assessment and Optimized visual-cue reasoning, yielding Bi-objective geo-Enhancement for the VLM in recognition and reasoning. GLOBE incorporates task-specific rewards that jointly enhance localizability assessment, visual-cue reasoning, and geolocation accuracy. Both qualitative and quantitative results demonstrate that GLOBE outperforms state-of-the-art open-source LVLMs on geo-localization tasks, particularly in diverse visual scenes, while also generating more insightful and interpretable reasoning trajectories. The data and code are available at https://github.com/lingli1996/GLOBE.

地理定位视觉语言模型推理增强可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。