构建多模态大模型地理定位评估基准,推动具身智能发展
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

- 设计三阶段观测场景:单视角、全景、具身移动,模拟真实感知过程
- 覆盖2207张全球街景图,测试感知、空间、常识与定位推理能力
- 揭示定位能力依赖综合认知,非单一视觉识别,适合研究具身智能者
多模态大语言模型在作为具身代理方面展现出强大潜力,但具身地理定位仍因缺乏细粒度评估而未被充分探索。我们提出ERGeoBench,一个面向视觉驱动的具身地理定位诊断基准。该基准在三个渐进设置下评估模型:单视角、全景视角和具身视角,其中代理可通过连续改变偏航、俯仰和缩放主动获取观测。基准包含2,207张分布全球的街景全景图,衡量四项互补能力:基础感知、空间意识、常识推理和地理定位推理。对领先专有及开源多模态大模型的评估显示,当前模型可推断高层次地理语义,但仍难以实现精细感知操作、度量定位以及跨视角的空间一致性。进一步观察发现,地理定位与其它能力维度高度相关,表明准确定位依赖于感知、空间推理与常识推断的整合,而非孤立的视觉识别。总体而言,ERGeoBench为诊断和推进类人具身地理定位提供了统一框架。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluation. We introduce ERGeoBench, a diagnostic benchmark for vision-driven embodied geo-localization. ERGeoBench evaluates models under three progressive settings -- single-view, panorama-view, and embodied-view -- where agents may actively acquire observations through sequential changes in yaw, pitch, and zoom. The benchmark contains 2,207 globally distributed street-view panoramas and measures four complementary capabilities: foundational perception, spatial awareness, common sense reasoning, and geo-localization reasoning. Evaluations of leading proprietary and open-source MLLMs show that current models can infer high-level geographic semantics, but still struggle with fine-grained perceptual operations, metric localization, and spatial consistency across views. We further observe that geo-localization is strongly correlated with the other capability dimensions, suggesting that accurate localization depends on integrated perception, spatial reasoning, and commonsense inference rather than isolated visual recognition. Overall, ERGeoBench provides a unified framework for diagnosing and advancing human-like embodied geo-localization. Project Page: https://kaixuewen.github.io/ERGeoBench/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。