评测大模型在遥感任务中使用工具的推理能力,填补领域空白。
ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
- 设计结构化工具交互流程,支持多步规划与遥感数据处理
- 覆盖486个任务、1778步专家验证推理,评估模型工具使用准确性
- 首次提供面向遥感领域的工具增强型模型评测基准,适合研究者与开发者
大语言模型(LLM)的进步使具备工具调用能力的智能体能够通过逐步推理解决复杂现实任务。然而,现有评估多集中于通用或跨模态场景,缺乏针对遥感等专业领域的工具使用评测基准。本文提出ThinkGeo,一个面向遥感任务的代理评测基准,通过结构化工具使用和多步规划,评估基于LLM的智能体表现。其包含由人类标注的多样化真实应用场景查询,涵盖城市规划、灾害评估、环境监测、交通分析等,数据来源为卫星或航拍影像,包括光学RGB和雷达(SAR)数据,要求智能体运用多种工具进行空间推理。采用类似ReAct的交互机制,在486个结构化任务上评估开放与闭源模型(如GPT-4o、Qwen2.5),共1,778步专家验证的推理步骤,报告每一步执行指标与最终答案正确率。分析显示不同模型在工具使用准确性和规划一致性方面存在显著差异。ThinkGeo提供了首个系统性测试平台,用于评估工具增强型LLM在遥感空间推理中的能力。
原文摘要 · Abstract (English)
Recent progress in large language models (LLMs) has enabled tool-augmented agents capable of solving complex real-world tasks through step-by-step reasoning. However, existing evaluations often focus on general-purpose or multimodal scenarios, leaving a gap in domain-specific benchmarks that assess tool-use capabilities in complex remote sensing use cases. We present ThinkGeo, an agentic benchmark designed to evaluate LLM-driven agents on remote sensing tasks via structured tool use and multi-step planning. Inspired by tool-interaction paradigms, ThinkGeo includes human-curated queries spanning a wide range of real-world applications such as urban planning, disaster assessment and change analysis, environmental monitoring, transportation analysis, aviation monitoring, recreational infrastructure, and industrial site analysis. Queries are grounded in satellite or aerial imagery, including both optical RGB and SAR data, and require agents to reason through a diverse toolset. We implement a ReAct-style interaction loop and evaluate both open and closed-source LLMs (e.g., GPT-4o, Qwen2.5) on 486 structured agentic tasks with 1,778 expert-verified reasoning steps. The benchmark reports both step-wise execution metrics and final answer correctness. Our analysis reveals notable disparities in tool accuracy and planning consistency across models. ThinkGeo provides the first extensive testbed for evaluating how tool-enabled LLMs handle spatial reasoning in remote sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。