让AI通过网页搜索和图像放大精准定位照片地点。
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
- 在推理中动态调用网页搜索和图像放大工具,实现多步地理定位。
- 在自建的GeoBench基准上表现超越多数开源模型,接近闭源大模型水平。
- 适合需要高精度地理推理的应用,如地图标注、旅游导航等。
当前的智能体视觉推理研究主要聚焦于图像编辑工具,缺乏面向通用场景的智能体模型。本文重新审视地理定位任务,该任务不仅需要精细的视觉定位,还需通过网络搜索验证或修正推理假设。由于现有基准无法满足高分辨率图像与深度推理挑战的需求,我们构建了GeoBench,包含全球范围内的照片与全景图,并附带多个城市的卫星图像子集,以严格评估智能体的地理定位能力。同时,我们提出GeoVista,一种将工具调用无缝集成到推理流程中的智能体模型,包括图像局部放大和网络搜索工具。我们设计了完整的训练流程:先进行冷启动监督微调以学习推理模式与工具使用先验,再通过强化学习进一步提升推理能力。采用分层奖励机制,利用多层级地理信息优化整体性能。实验表明,GeoVista在地理定位任务上显著优于其他开源模型,在多数指标上达到与Gemini-2.5-flash和GPT-5相当的水平。
原文摘要 · Abstract (English)
Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agentic models. In this work, we revisit the geolocalization task, which requires not only nuanced visual grounding but also web search to confirm or refine hypotheses during reasoning. Since existing geolocalization benchmarks fail to meet the need for high-resolution imagery and the localization challenge for deep agentic reasoning, we curate GeoBench, a benchmark that includes photos and panoramas from around the world, along with a subset of satellite images of different cities to rigorously evaluate the geolocalization ability of agentic models. We also propose GeoVista, an agentic model that seamlessly integrates tool invocation within the reasoning loop, including an image-zoom-in tool to magnify regions of interest and a web-search tool to retrieve related web information. We develop a complete training pipeline for it, including a cold-start supervised fine-tuning (SFT) stage to learn reasoning patterns and tool-use priors, followed by a reinforcement learning (RL) stage to further enhance reasoning ability. We adopt a hierarchical reward to leverage multi-level geographical information and improve overall geolocalization performance. Experimental results show that GeoVista surpasses other open-source agentic models on the geolocalization task greatly and achieves performance comparable to closed-source models such as Gemini-2.5-flash and GPT-5 on most metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。