用分层智能体+外部工具验证,提升图像地理定位准确性
LocationAgent: A Hierarchical Agent for Image Geolocation via Decoupling Strategy and Evidence from Parametric Knowledge
- 分层设计推理-执行-记录架构,防止多步推理偏差
- 零样本场景下性能领先现有方法至少30%
- 专为中文数据稀缺问题构建中国城市定位基准测试集
图像地理定位旨在基于视觉内容推断拍摄位置。该任务本质上是包含假设-验证循环的推理过程,要求模型兼具地理空间推理能力与证据验证能力。现有方法通常通过监督训练或轨迹强化微调将位置知识和推理模式固化于静态记忆中,导致在开放世界或需动态知识的场景下易产生事实幻觉和泛化瓶颈。为此,我们提出分层定位智能体LocationAgent,核心思想是在模型中保留分层推理逻辑的同时,将地理证据验证任务交由外部工具完成。为实现分层推理,我们设计了RER架构(Reasoner-Executor-Recorder),通过角色分离与上下文压缩避免多步推理漂移;为支持证据验证,构建了一套线索探索工具,提供多样证据辅助定位推理。此外,针对现有数据集存在数据泄露及中文数据匮乏的问题,我们引入CCL-Bench(中国城市定位基准测试集),涵盖多种场景粒度与难度等级。大量实验表明,LocationAgent在零样本设置下性能显著优于现有方法,最低提升达30%。
原文摘要 · Abstract (English)
Image geolocation aims to infer capture locations based on visual content. Fundamentally, this constitutes a reasoning process composed of \textit{hypothesis-verification cycles}, requiring models to possess both geospatial reasoning capabilities and the ability to verify evidence against geographic facts. Existing methods typically internalize location knowledge and reasoning patterns into static memory via supervised training or trajectory-based reinforcement fine-tuning. Consequently, these methods are prone to factual hallucinations and generalization bottlenecks in open-world settings or scenarios requiring dynamic knowledge. To address these challenges, we propose a Hierarchical Localization Agent, called LocationAgent. Our core philosophy is to retain hierarchical reasoning logic within the model while offloading the verification of geographic evidence to external tools. To implement hierarchical reasoning, we design the RER architecture (Reasoner-Executor-Recorder), which employs role separation and context compression to prevent the drifting problem in multi-step reasoning. For evidence verification, we construct a suite of clue exploration tools that provide diverse evidence to support location reasoning. Furthermore, to address data leakage and the scarcity of Chinese data in existing datasets, we introduce CCL-Bench (China City Location Bench), an image geolocation benchmark encompassing various scene granularities and difficulty levels. Extensive experiments demonstrate that LocationAgent significantly outperforms existing methods by at least 30\% in zero-shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。