arXiv:2608.03882cs.CLcs.AI2026-08

测试大模型地理推理能力,覆盖全球201国多语言数据

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

论文配图:MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning
图 1 · 摘自论文原文
  • 构建跨语言、全球覆盖的地理推理基准,含4.6万题对
  • 模型在网格索引与形状计算上表现差,即使给正确知识也难提升
  • 低收入地区问题模型更弱,暴露出系统性偏差

地理推理(计算实体间距离、包含关系等空间关系)对导航与物流至关重要,但大语言模型虽存储大量地理知识,仍难以完成所需几何与拓扑计算。现有基准存在合成性、规模小、单一语言和地理覆盖有限等问题。本文提出MultiGlobeQA,一个包含46,060个问答对的多语言基准,涵盖14类空间功能与15种答案格式,基于三个知识图谱的执行级真值。通过收入与密度分层采样覆盖201个国家和地区,支持英语及16种高/低资源语言的并行问题。在参数化、推理与代理设置下,模型在网格索引和形状计算任务上表现严重退化,拓扑关系与方向判断相对较好。检索与工具使用带来显著提升,但即便提供真实事实,性能仍低于三分之二,表明瓶颈在于计算而非知识获取。模型在低收入地区表现更差,且真实事实反而扩大该差距。

原文摘要 · Abstract (English)

Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computation despite storing considerable geographic knowledge. Existing benchmarks localize these failures only partially: they are synthetic or smallscale, largely monolingual, and offer limited control over geographic coverage. We introduce MultiGlobeQA, a multilingual benchmark of 46,060 question-answer pairs spanning 14 spatial-function families and 15 answer formats, with execution-based ground truth over three knowledge graphs. It covers 201 countries and territories via income- and density-stratified sampling, with parallel questions in English and 16 additional high- and low-resource languages. Across parametric, reasoning, and agentic settings, LLMs collapse on tasks requiring grid indexing and shape computation, while topological relations and directions fare best. Retrieval and tool use yield considerable gains, yet performance plateaus below two thirds even when gold facts are supplied, indicating that computation, not access to knowledge, is the bottleneck. Models also underperform on low-income regions, a gap that gold facts widen rather than close.

地理推理多语言大模型评估偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。