arXiv:2509.18173cs.LGcs.CL2025-09EMNLP

构建首个地理路线逆向认知基准,评估大模型理解空间路径能力。

TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route

  • 设计双向转换工具PathBuilder,打通语言与导航路线的语义桥梁。
  • 在3.6万条全球城市路线数据上测试,多数模型无法准确逆向返回起点。
  • 揭示大模型在路线生成中信心过强、鲁棒性差的核心缺陷,适合评估者参考。

人类能通过自然语言理解地理空间信息,而大语言模型(LLMs)在该领域的认知能力尚未充分探索。以往研究受限于不可量化的指标、有限的评测数据集和模糊的研究层次。为此,我们提出一个大规模基准,并对LLMs的地理路线认知能力进行全面评估。构建了包含来自全球12个大城市的36,000条路线的大规模评测数据集。提出PathBuilder,一种可将自然语言指令转换为导航路线、反之亦然的新工具,弥合地理信息与自然语言之间的鸿沟。设计新的评估框架与度量标准,严格评估11个前沿(SOTA)LLMs在路线逆向任务中的表现。结果显示,LLMs在路线逆向方面存在显著局限:多数逆向路径既无法返回起点,也与最优路径不相似。此外,模型在路线生成中表现出低鲁棒性,且对错误答案具有高置信度。代码与数据已公开于:TurnBack。

原文摘要 · Abstract (English)

Humans can interpret geospatial information through natural language, while the geospatial cognition capabilities of Large Language Models (LLMs) remain underexplored. Prior research in this domain has been constrained by non-quantifiable metrics, limited evaluation datasets and unclear research hierarchies. Therefore, we propose a large-scale benchmark and conduct a comprehensive evaluation of the geospatial route cognition of LLMs. We create a large-scale evaluation dataset comprised of 36000 routes from 12 metropolises worldwide. Then, we introduce PathBuilder, a novel tool for converting natural language instructions into navigation routes, and vice versa, bridging the gap between geospatial information and natural language. Finally, we propose a new evaluation framework and metrics to rigorously assess 11 state-of-the-art (SOTA) LLMs on the task of route reversal. The benchmark reveals that LLMs exhibit limitation to reverse routes: most reverse routes neither return to the starting point nor are similar to the optimal route. Additionally, LLMs face challenges such as low robustness in route generation and high confidence for their incorrect answers. Code\ \&\ Data available here: \href{https://github.com/bghjmn32/EMNLP2025_Turnback}{TurnBack.}

地理认知大模型评估路线逆向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。