arXiv:2606.12821cs.AIcs.ET2026-06

首个面向环境地理分析的AI代理评测基准,验证大模型真实调用地理接口的能力。

GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models

论文配图:GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models
图 1 · 摘自论文原文
  • 构建真实地理API接口上的结构化工具调用评测体系,覆盖93项任务
  • Claude Sonnet 4准确率达60.8%,开源模型DeepSeek V3.2成本仅为1/11
  • 揭示对比推理等任务普遍失败,凸显当前模型在复杂地理决策中的局限

环境科学家将大量时间用于数据处理而非分析,而能够自动化地理空间工作流的AI代理尚未经过有效验证:缺乏基于真实API的结构化工具调用评测基准。本文提出GeoNatureAgent Benchmark,首个通过结构化工具调用真实生产级地理空间API的环境分析代理评测基准。包含93个任务,涵盖18类场景,包括市镇分析、多轮对话、空间推理、跨指标整合、错误处理与恢复、排序、比较、多语言理解、生境分析及任务拒绝。任务基于一个开放可自托管的API进行评估,该接口提供西班牙和葡萄牙的三项环境指标,共16个工具。我们评估了七种LLM(Claude Sonnet 4、DeepSeek V3.2、GLM-5、Gemini 2.5 Pro、Qwen3-235B、GPT-OSS-120B、Llama 4 Scout)在三个温度为1.0的随机种子下的表现,报告能力与单例成本作为两个独立维度。结果显示:(1) Claude Sonnet 4以60.8%±0.8%领先,DeepSeek V3.2为56.3%±3.1%,其余模型均未超过51%;(2) 成本-准确率帕累托前沿主要由开源模型占据,DeepSeek V3.2实现93%的性能,成本仅为Claude的1/11($0.011/案例);(3) 比较任务普遍失败(接近值比较准确率为0%),暴露系统性推理缺陷;(4) 与通用GIS基准相比,真实接口调用更具区分度,准确率低25–35个百分点。我们进一步通过集成BigEarthNet V2土地覆盖数据扩展至葡萄牙区域。基准、评测套件与自托管API已公开。

原文摘要 · Abstract (English)

Environmental scientists spend disproportionate effort on data wrangling rather than analysis, and AI agents that automate geospatial workflows remain unvalidated: no benchmark evaluates agents operating through structured tool calling against real APIs. We introduce the GeoNatureAgent Benchmark, the first benchmark for environmental analysis agents that operate via structured tool calls to a production-style geospatial API. It comprises 93 tasks across 18 categories, covering municipality analysis, multi-turn conversation, spatial reasoning, cross-indicator synthesis, error handling and recovery, ranking, comparison, multilingual understanding, habitat analysis, and task rejection. Tasks are evaluated against an open, self-hostable API serving three environmental indicators across Spain and Portugal via sixteen tools. We evaluate seven LLMs (Claude Sonnet 4, DeepSeek V3.2, GLM-5, Gemini 2.5 Pro, Qwen3-235B, GPT-OSS-120B, Llama 4 Scout) under three temperature-1.0 seeds, reporting capability and per-case cost as orthogonal axes. We find: (1) Claude Sonnet 4 leads at 60.8% +/- 0.8%, followed by DeepSeek V3.2 at 56.3% +/- 3.1%, with no other model above 51%; (2) the cost-accuracy Pareto frontier is occupied mostly by open-weight models, with DeepSeek V3.2 offering 93% of Claude's capability at 11x lower cost ($0.011/case); (3) comparison tasks remain universally unsolved (0% on close-value comparisons), exposing systematic reasoning limits; and (4) structured tool calling against a real API is more discriminative than general-purpose GIS benchmarks, with accuracies 25-35 points lower. We further show extensibility by integrating BigEarthNet V2 land cover for Portugal alongside Spanish CO2 and erosion indicators. The benchmark, harness, and self-hostable API are publicly available.

地理分析AI代理大模型评测开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。