提出自适应推理框架,让图像定位更准更少出错。
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
- 用可定位性评分动态调整推理深度,避免无效计算。
- 在多个数据集上达最新水平,错误率显著降低。
- 适合需要高精度定位的地理图像应用。
视觉语言模型(VLMs)为全局图像地理定位带来了新范式,通过检索增强生成(RAG)和基于推理的推断。然而,RAG方法受限于检索数据库质量,而基于推理的方法无法内化图像可定位性,依赖低效且固定的推理路径,导致幻觉增多、准确率下降。为此,我们提出优化的可定位性得分,量化图像在地理定位中进行深度推理的适配度。基于该指标,构建了包含复杂视觉场景增强推理轨迹的Geo-ADAPT-51K数据集。在此基础上,提出两阶段组相对策略优化(GRPO)课程,设计定制化奖励函数以调控推理深度、视觉定位和层级地理准确性。所提出的Geo-ADAPT框架学习自适应推理策略,在多个地理定位基准上取得最先进性能,并显著减少幻觉。
原文摘要 · Abstract (English)
The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven inference. However, RAG methods are constrained by retrieval database quality, while reasoning-driven approaches fail to internalize image locatability, relying on inefficient, fixed-depth reasoning paths that increase hallucinations and degrade accuracy. To overcome these limitations, we introduce an Optimized Locatability Score that quantifies an image's suitability for deep reasoning in geo-localization. Using this metric, we curate Geo-ADAPT-51K, a locatability-stratified reasoning dataset enriched with augmented reasoning trajectories for complex visual scenes. Building on this foundation, we propose a two-stage Group Relative Policy Optimization (GRPO) curriculum with customized reward functions that regulate adaptive reasoning depth, visual grounding, and hierarchical geographical accuracy. Our framework, Geo-ADAPT, learns an adaptive reasoning policy, achieves state-of-the-art performance across multiple geo-localization benchmarks, and substantially reduces hallucinations by reasoning both adaptively and efficiently.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。