arXiv:2608.01645cs.AI2026-08

构建首个真实地理分析任务基准,评估大模型在复杂空间工作流中的表现。

GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks

论文配图:GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks
图 1 · 摘自论文原文
  • 从专业论坛收集349个真实地理任务,基于真实数据生成可执行流程
  • 每项任务配有精确真值输出,支持严格匹配而非依赖模型主观判断
  • 测试显示最佳模型仅完成32.7%任务,凸显大模型处理复杂地理分析的困难

地理信息系统(GIS)专业人士依赖多步骤空间分析流程支持城市规划、灾害响应和环境监测决策。该过程繁琐、耗时且易出错。尽管具备外部工具的大语言模型(LLM)代理有望自动化地理空间分析,但其在真实GIS工作流中的能力仍缺乏探索。现有基准多源自教材或模型生成,规模小、流程浅,且无真值输出,依赖代码相似性或模型评判,易混淆流程相似与任务正确。为此,我们提出GISAgentBench,一个由349个多步GIS任务组成的基准,数据源自GIS Stack Exchange,覆盖六个实际地理区域,均基于真实公开数据。每个任务配备可执行参考流程与精确真值输出文件,实现严格、确定性的容差感知输出匹配。对六种LLM模型的评估表明,真实地理工作流仍具挑战:最佳代理在严格容差评分下仅完成32.7%的任务,尽管多数模型输出接近真值。

原文摘要 · Abstract (English)

Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone. While recent large language model (LLM) agents equipped with external tools have the potential to automate geospatial analysis, their ability to perform realistic GIS workflows remains largely unexplored. Existing GIS agent benchmarking datasets are mostly drawn from textbooks, tutorials, or LLM-generated seeds and remain limited in size and trajectory depth. More importantly, none provides ground truth outputs. They therefore rely on surrogate signals such as code similarity, trajectory matching, or LLM and VLM judges, which can conflate workflow resemblance with task correctness. To address this gap, we introduce GISAgentBench, a benchmark of 349 multi-step GIS tasks curated from GIS Stack Exchange and instantiated on real public data across six selected geographic areas of interest. Each task ships with an executable reference trajectory and an exact ground truth output file, enabling strict, deterministic, tolerance-aware output matching beyond LLM judging. Evaluations of six LLM models reveal that realistic GIS workflows remain challenging: the best agent completes only 32.7% of tasks under strict tolerance-aware scoring, although most models produce outputs that are close to the ground truth.

GIS大模型代理基准测试空间分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。