arXiv:2606.13148cs.AI2026-06被引 2

构建首个可执行的地球系统多源数据推理基准,让AI能像科学家一样处理复杂环境任务。

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

论文配图:TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?
图 1 · 摘自论文原文
  • 用ReAct框架将语言推理与科学工具联动,实现跨数据类型交互
  • 包含403个任务、24,500步验证执行,覆盖8大应用领域
  • 首次引入过程级工具使用评估与容差数值评分,适合气候与地理科研人员

气候与环境决策日益需要对异构数据进行综合推理,包括网格化物理数据、卫星影像、地理空间上下文和模拟器输出。气象与气候基础模型预测能力较强,但无法进行语言交互式推理;而大语言模型(LLMs)虽能进行语言推理,却不能直接操作高维地球系统数据。因此,真实地球科学工作流仍缺乏支持。我们提出TerraBench,一个基于TerraAgent的可执行地球科学推理基准,该框架采用ReAct风格,通过推理、工具调用与观测的交替执行,将LLM规划与环境检索、地理空间处理、模拟及基于成果的计算工具耦合。TerraBench统一了地球观测影像分析、网格数据处理、GIS推理与模拟,在单一可执行界面中完成,而以往基准将其拆分为孤立任务。它也是该领域首个将过程级工具使用度量与容差感知数值评分结合的基准。该基准包含403个综合性智能体任务,分为三个赛道(基础、模拟器驱动、文档验证),覆盖八个应用领域,共24,500个已验证执行步骤。结果表明,可靠的地球科学智能体必须超越简单工具访问,实现异构工作流协调、精确参数化工具,并保持成果溯源性。

原文摘要 · Abstract (English)

Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate foundation models can forecast well, but do not reason interactively in language, while large language models (LLMs) reason in language but cannot operate directly on high-dimensional Earth-system data. As a result, real scientific workflows in Earth-science remain underserved. We introduce TerraBench, a benchmark for grounded Earth-science reasoning, built on TerraAgent, a ReAct-style executable framework that interleaves reasoning, tool calls, and observations to couple LLM planning with scientific tools for environmental retrieval, geospatial processing, simulation, and artifact-backed computation. TerraBench unifies analysis of Earth observation imagery, gridded data, GIS reasoning and simulation in a single executable interface, whereas prior benchmarks isolate these capabilities into narrow individual tasks. It is also the first in this space to pair process-level tool-use metrics with tolerance-aware numeric scoring. The benchmark comprises 403 extensive agentic tasks across three tracks (Fundamentals, Simulator-Grounded, and Document-Grounded Verification) and eight application domains with 24,500 verified execution steps. These results indicate that reliable Earth-science agents must go beyond tool access to coordinate heterogeneous workflows, parameterize tools precisely, and preserve artifact provenance.

地球科学智能体多模态推理可执行基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。