arXiv:2604.13888cs.AI2026-04被引 1

构建动态地理分析评估基准,提升大模型在空间任务中的执行准确率。

GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis

论文配图:GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
图 1 · 摘自论文原文
  • 设计动态交互式评估环境,集成117个原子地理工具。
  • 提出参数执行准确率(PEA)指标,精准衡量参数推断效果。
  • 引入计划-反应架构,显著提升复杂任务的容错与推理能力。

将大语言模型(LLMs)融入地理信息系统(GIS)标志着自主空间分析范式的转变。然而,由于地理空间工作流具有多步骤、复杂的特性,评估基于LLM的智能体仍面临挑战。现有基准主要依赖静态文本或代码匹配,忽略了动态运行时反馈及空间输出的多模态特性。为此,我们提出GeoAgentBench(GABench),一个面向工具增强型GIS智能体的动态交互评估基准。GABench提供真实执行沙箱,集成117个原子GIS工具,覆盖6个核心GIS领域中的53项典型空间分析任务。鉴于精确参数配置是动态GIS环境中执行成功的关键,我们设计了参数执行准确率(PEA)指标,采用“最后尝试对齐”策略量化隐式参数推断的保真度。同时,引入基于视觉-语言模型(VLM)的验证机制,评估数据空间准确性与制图风格一致性。为应对因参数不匹配和运行时异常导致的频繁任务失败,我们开发了新型智能体架构Plan-and-React,通过解耦全局规划与逐步反应执行,模拟专家认知流程。七种代表性LLM的大量实验表明,Plan-and-React范式显著优于传统框架,在多步推理与错误恢复方面实现逻辑严谨性与执行鲁棒性的最佳平衡。研究揭示了当前能力边界,并建立了下一代自主地理人工智能评估的坚实标准。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) into Geographic Information Systems (GIS) marks a paradigm shift toward autonomous spatial analysis. However, evaluating these LLM-based agents remains challenging due to the complex, multi-step nature of geospatial workflows. Existing benchmarks primarily rely on static text or code matching, neglecting dynamic runtime feedback and the multimodal nature of spatial outputs. To address this gap, we introduce GeoAgentBench (GABench), a dynamic and interactive evaluation benchmark tailored for tool-augmented GIS agents. GABench provides a realistic execution sandbox integrating 117 atomic GIS tools, encompassing 53 typical spatial analysis tasks across 6 core GIS domains. Recognizing that precise parameter configuration is the primary determinant of execution success in dynamic GIS environments, we designed the Parameter Execution Accuracy (PEA) metric, which utilizes a "Last-Attempt Alignment" strategy to quantify the fidelity of implicit parameter inference. Complementing this, a Vision-Language Model (VLM) based verification is proposed to assess data-spatial accuracy and cartographic style adherence. Furthermore, to address the frequent task failures caused by parameter misalignments and runtime anomalies, we developed a novel agent architecture, Plan-and-React, that mimics expert cognitive workflows by decoupling global orchestration from step-wise reactive execution. Extensive experiments with seven representative LLMs demonstrate that the Plan-and-React paradigm significantly outperforms traditional frameworks, achieving the optimal balance between logical rigor and execution robustness, particularly in multi-step reasoning and error recovery. Our findings highlight current capability boundaries and establish a robust standard for assessing and advancing the next generation of autonomous GeoAI.

地理智能大模型评估空间分析智能体架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。