arXiv:2608.29483cs.CVcs.CL2026-08中稿 · EMNLP

用沉浸式导航评估视觉语言模型的地理定位能力

GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

论文配图:GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation
图 1 · 摘自论文原文
  • 构建可自主导航的地理定位环境,让模型通过移动观察提升判断
  • 静态图像模型准确率高但难识区域特征,导航后整体精度显著提升
  • 发现模型对发达/发展中地区存在严重偏见,且纠错能力差

现代视觉语言模型在图像地理定位任务上已超越人类基准,该任务在灾难响应、开源情报验证和位置隐私中至关重要。然而,现有研究多局限于静态图像检索、分类与预测。我们主张,真实任务应包含具身导航——多模态代理自主探索环境,收集观测后再作判断。为此,我们提出GeoAgent,一个基于代理的基准测试环境,要求代理在街景环境中导航,通过序列推理优化地理定位。分析显示,现代VLM虽能在国家和大洲级别取得良好表现,却难以识别区域模式。相比静态图像基线,具身导航显著提升各项指标准确性。同时发现,前沿模型在发达与发展中国家间存在严重偏差,且在错误先验下自我改进能力极弱。本工作揭示了具身导航与空间推理的挑战。代码与环境已公开:https://geoagent-benchmark.github.io

原文摘要 · Abstract (English)

Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response, OSINT verification, and location privacy. However, most efforts to study AI behavior on the task remain limited to static image-based retrieval, classification, and predictions. We argue that faithful recreation of the task should involve embodied navigation, where a multimodal agent autonomously explores its surroundings to gather observations before submitting a prediction. To this end, we introduce \textbf{GeoAgent}, an agentic environment-based benchmark that requires agents to navigate Street View environments to refine their geolocalization through sequential reasoning. Our analysis shows that modern VLMs struggle to discern regional patterns while succeeding at country- and continent-level predictions. When compared to static image-based baselines, agentic navigation significantly improves accuracy across established metrics. We also note severe bias in a developed/developing region context across frontier model architectures and poor self-improvement capabilities given incorrect priors. Overall, our work establishes the challenges of embodied navigation and geospatial reasoning. We publicly release our code and the GeoAgent environment: https://geoagent-benchmark.github.io

视觉语言模型地理定位具身智能导航基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。