GeoAgent通过地理特征强化学习,精准定位地址并生成人类级推理过程。
GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics
- 基于地理专家标注的思维链数据,设计地理相似性与一致性奖励机制。
- 在多粒度任务中超越现有方法与通用视觉语言模型表现。
- 适合需要高精度地理推理与可解释性的应用,如地图服务、导航系统。
本文提出GeoAgent,一种能与人类紧密协作并推导细粒度地址结论的模型。以往基于强化学习的方法虽在性能和可解释性上取得突破,但仍依赖AI生成的思维链(CoT)数据与训练策略,与地理特性存在冲突。为此,我们首先构建了由地理专家和专业玩家标注的新型地理定位数据集GeoSeek。进一步深入探索地理任务内在特征,提出由一致性代理评估的地理相似性奖励与一致性奖励,引导模型从地理视角收敛至正确答案,同时保障推理过程的完整与一致。实验表明,GeoAgent在多个粒度下均优于现有方法及一系列通用视觉语言模型,且生成的推理过程高度贴近人类思维。
原文摘要 · Abstract (English)
This paper presents GeoAgent, a model capable of reasoning closely with humans and deriving fine-grained address conclusions. Previous RL-based methods have achieved breakthroughs in performance and interpretability but still remain concerns because of their reliance on AI-generated chain-of-thought (CoT) data and training strategies, which conflict with geographic characteristics. To address these issues, we first introduce GeoSeek, a new geolocation dataset comprising CoT data annotated by geographic experts and professional players. We further thoroughly explore the inherent characteristics of geographic tasks and propose a geo-similarity reward and a consistency reward assessed by a consistency agent to assist training. This encourages the model to converge towards correct answers from a geographic perspective while ensuring the integrity and consistency of its reasoning process. Experimental results show that GeoAgent outperforms existing methods and a series of general VLLMs across multiple grains, while generating reasoning that closely aligns with humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。