arXiv:2503.18129cs.CLcs.AI2025-03被引 18

评测大模型在复杂地理任务中的工具调用能力,发现o4-mini和Claude表现最佳。

GeoBenchX: Benchmarking LLMs in Agent Solving Multistep Geospatial Tasks

  • 构建包含23个地理函数的工具调用代理,测试多步骤地理空间任务
  • o4-mini和Claude 3.5 Sonnet综合表现最优,部分模型误判不可解任务
  • 开源评估框架与数据集,助力地理AI大模型标准化评测

本文建立了一个针对大型语言模型(LLMs)在多步骤地理空间任务中工具调用能力的基准测试。评估了八款商用大模型(Claude Sonnet 3.5和4、Claude Haiku 3.5、Gemini 2.0 Flash、Gemini 2.5 Pro Preview、GPT-4o、GPT-4.1和o4-mini),使用配备23个地理空间功能的简单工具调用代理。基准涵盖四类递增复杂度的任务,包含可解与故意不可解任务以检验拒绝准确性。开发了基于LLM作为评判者的评估框架,对比代理解决方案与参考答案。结果显示,o4-mini和Claude 3.5 Sonnet整体表现最佳;OpenAI的GPT-4.1、GPT-4o及Google的Gemini 2.5 Pro Preview紧随其后,但后者在识别不可解任务上更高效。Claude Sonnet 4因倾向于提供任何解而非拒绝任务,准确率较低。观察到显著的令牌消耗差异,Anthropic模型的耗能高于竞争对手。常见错误包括误解几何关系、依赖过时知识、数据操作效率低下。所生成的基准数据集、评估框架与数据生成流水线已开源(地址:https://github.com/Solirinai/GeoBenchX),为地理人工智能大模型的持续评估提供标准化方法。

原文摘要 · Abstract (English)

This paper establishes a benchmark for evaluating tool-calling capabilities of large language models (LLMs) on multi-step geospatial tasks relevant to commercial GIS practitioners. We assess eight commercial LLMs (Claude Sonnet 3.5 and 4, Claude Haiku 3.5, Gemini 2.0 Flash, Gemini 2.5 Pro Preview, GPT-4o, GPT-4.1 and o4-mini) using a simple tool-calling agent equipped with 23 geospatial functions. Our benchmark comprises tasks in four categories of increasing complexity, with both solvable and intentionally unsolvable tasks to test rejection accuracy. We develop a LLM-as-Judge evaluation framework to compare agent solutions against reference solutions. Results show o4-mini and Claude 3.5 Sonnet achieve the best overall performance, OpenAI's GPT-4.1, GPT-4o and Google's Gemini 2.5 Pro Preview do not fall far behind, but the last two are more efficient in identifying unsolvable tasks. Claude Sonnet 4, due its preference to provide any solution rather than reject a task, proved to be less accurate. We observe significant differences in token usage, with Anthropic models consuming more tokens than competitors. Common errors include misunderstanding geometrical relationships, relying on outdated knowledge, and inefficient data manipulation. The resulting benchmark set, evaluation framework, and data generation pipeline are released as open-source resources (available at https://github.com/Solirinai/GeoBenchX), providing one more standardized method for the ongoing evaluation of LLMs for GeoAI.

地理智能大模型评测工具调用开源基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。