arXiv:2509.05881cs.SEcs.AI2025-09被引 28

构建首个地理空间分析基准测试,评估大模型生成代码与工作流的能力

GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation

  • 基于50个真实地理任务设计可验证的评测框架
  • 闭源模型代码正确率95%,开源模型仅48.5%有效
  • 适用于地理信息科学自动化研究者与模型开发者

大语言模型在地理空间分析和GIS工作流自动化中的应用日益受到关注,但其实际能力仍不明确。为此,我们提出GeoAnalystBench,一个包含50个基于Python的地理处理任务的基准测试,所有任务均来自真实地理问题并经地理信息专家严格验证。每个任务配有最小交付成果,评估涵盖工作流有效性、结构对齐度、语义相似性及代码质量(CodeBLEU)。实验评估了多个专有与开源模型。结果显示:如ChatGPT-4o-mini等闭源模型在有效性上达95%,代码对齐度高(CodeBLEU 0.39);而DeepSeek-R1-7B等小型开源模型则常生成不完整或不一致的工作流(有效性48.5%,CodeBLEU 0.272)。涉及深层空间推理的任务(如空间关系识别、最优选址)仍是所有模型的难点。该研究揭示了当前大模型在GIS自动化中的潜力与局限,并提供可复现的人机协同框架以推动地理人工智能研究。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have fueled growing interest in automating geospatial analysis and GIS workflows, yet their actual capabilities remain uncertain. In this work, we call for rigorous evaluation of LLMs on well-defined geoprocessing tasks before making claims about full GIS automation. To this end, we present GeoAnalystBench, a benchmark of 50 Python-based tasks derived from real-world geospatial problems and carefully validated by GIS experts. Each task is paired with a minimum deliverable product, and evaluation covers workflow validity, structural alignment, semantic similarity, and code quality (CodeBLEU). Using this benchmark, we assess both proprietary and open source models. Results reveal a clear gap: proprietary models such as ChatGPT-4o-mini achieve high validity 95% and stronger code alignment (CodeBLEU 0.39), while smaller open source models like DeepSeek-R1-7B often generate incomplete or inconsistent workflows (48.5% validity, 0.272 CodeBLEU). Tasks requiring deeper spatial reasoning, such as spatial relationship detection or optimal site selection, remain the most challenging across all models. These findings demonstrate both the promise and limitations of current LLMs in GIS automation and provide a reproducible framework to advance GeoAI research with human-in-the-loop support.

地理信息大模型评估代码生成空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。