让AI像人一样探索图像找位置,提升全球定位精度。
Learning to Wander: Improving the Global Image Geolocation Ability of LMMs via Actionable Reasoning
- 用可行动的推理计划替代纯文字思考,主动调整视角和位置。
- 在32000张全景图上测试,定位精度显著优于现有模型。
- 适合研究具身智能、视觉导航与地理信息融合的学者。
地理定位任务需要丰富的世界知识和复杂推理能力。尽管先进大型多模态模型(LMMs)具备上述能力,但其在地理定位任务上的表现仍不明确。为此,我们提出首个开放获取的全球地理定位基准——WanderBench,包含覆盖六大洲超过32,000张全景图,构建为可导航图结构,支持旋转、移动等物理动作,将地理定位从静态识别转变为交互式探索。基于此,我们提出GeoAoT(Action of Thought)框架,将推理与具身动作结合,不生成文本推理链,而是输出如靠近地标或调整视角等可执行计划,主动降低不确定性。我们还设计了联合评估协议,同时衡量定位准确率与难易感知提问能力。在19个大型多模态模型上的实验表明,GeoAoT在细粒度定位和动态环境泛化方面表现更优。WanderBench与GeoAoT定义了具身视觉理解中以推理驱动的可行动地理定位新范式。
原文摘要 · Abstract (English)
Geolocation, the task of identifying the geographic location of an image, requires abundant world knowledge and complex reasoning abilities. Though advanced large multimodal models (LMMs) have shown superior aforementioned capabilities, their performance on the geolocation task remains unexplored. To this end, we introduce \textbf{WanderBench}, the first open access global geolocation benchmark designed for actionable geolocation reasoning in embodied scenarios. WanderBench contains over 32K panoramas across six continents, organized as navigable graphs that enable physical actions such as rotation and movement, transforming geolocation from static recognition into interactive exploration. Building on this foundation, we propose \textbf{GeoAoT} (Action of Thought), a \underline{Geo}location framework with \underline{A}ction of \underline{T}hough, which couples reasoning with embodied actions. Instead of generating textual reasoning chains, GeoAoT produces actionable plans such as, approaching landmarks or adjusting viewpoints, to actively reduce uncertainty. We further establish an evaluation protocol that jointly measures geolocation accuracy and difficulty-aware geolocation questioning ability. Experiments on 19 large multimodal models show that GeoAoT achieves superior fine-grained localization and stronger generalization in dynamic environments. WanderBench and GeoAoT define a new paradigm for actionable, reasoning driven geolocation in embodied visual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。