用智能体推理让大模型精准定位,还能验证结果真伪。
SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning
- 构建智能体流程,结合工具搜索与地图验证视觉线索。
- 在多个基准上达到顶尖表现,显著减少错误定位。
- 适合需要可靠地理定位的自动驾驶与地图应用。
大型视觉语言模型在地理定位任务中展现强大推理能力,但在真实场景下常因视觉线索稀疏、分布不均且高度模糊而表现不佳。以往方法受限于内部知识,难以提供可验证结果,面对混淆证据时仍给出自信但无根据的预测。为此,我们提出SpotAgent,将地理定位建模为一种智能体推理过程,通过专家级推理融合视觉理解与工具辅助验证。SpotAgent利用外部工具(如网络搜索、地图)主动探索并验证视觉线索,采用ReAct框架实现。我们设计了三阶段后训练流程:首先进行监督微调以实现基础对齐;接着通过多智能体框架生成高质量轨迹,开展智能体冷启动,培养工具调用能力;最后通过强化学习优化推理能力。提出空间感知动态筛选策略,在强化学习阶段优先选择空间难度高的样本以提升效率。大量实验表明,SpotAgent在标准基准上达到当前最优性能,有效缓解幻觉问题,实现精确且可验证的地理定位。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have demonstrated strong reasoning capabilities in geo-localization, yet they often struggle in real-world scenarios where visual cues are sparse, long-tailed, and highly ambiguous. Previous approaches, bound by internal knowledge, often fail to provide verifiable results, yielding confident but ungrounded predictions when faced with confounded evidence. To address these challenges, we propose SpotAgent, a framework that formalizes geo-localization into an agentic reasoning process that leverages expert-level reasoning to synergize visual interpretation with tool-assisted verification. SpotAgent actively explores and verifies visual cues by leveraging external tools (e.g., web search, maps) through a ReAct diagram. We introduce a 3-stage post-training pipeline starting with a Supervised Fine-Tuning (SFT) stage for basic alignment, followed by an Agentic Cold Start phase utilizing high-quality trajectories synthesized via a Multi-Agent framework, aiming to instill tool-calling expertise. Subsequently, the model's reasoning capabilities are refined through Reinforcement Learning. We propose a Spatially-Aware Dynamic Filtering strategy to enhance the efficiency of the RL stage by prioritizing learnable samples based on spatial difficulty. Extensive experiments on standard benchmarks demonstrate that SpotAgent achieves state-of-the-art performance, effectively mitigating hallucinations while delivering precise and verifiable geo-localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。