构建首个大气科学大模型评测基准,系统评估模型在五大领域的推理能力。
AtmosSci-Bench: Evaluating the Recent Advance of Large Language Model for Atmospheric Science
- 设计多选与开放题双格式评测框架,覆盖五大学科方向。
- 测试四类模型,发现数学增强型模型在物理推导中表现更优。
- 适合气候研究者、大模型开发者及科学计算领域应用者参考。
大语言模型(LLM)在推理能力上的快速进步,为解决大气科学中的复杂问题并推动科学发现带来变革性潜力。然而,有效利用LLM需一个全面且稳健的评估基准。为此,我们提出AtmosSci-Bench,一个针对大气科学五大核心问题领域的新型评测基准:水文、大气动力学、大气物理学、地球物理学和物理海洋学。该基准采用双格式设计,包含多选题(MCQs)与开放题(OEQs),支持自动化评估与深层概念理解分析。通过模板化生成框架创建多样化、研究生水平的问题,并引入符号扰动以增强多样性;开放题用于考察开放式推理能力。我们对代表性模型进行了全面评估,分为四类:指令微调模型、高级推理模型、数学增强模型和领域专用气候模型。分析揭示了各类模型在大气科学任务中的推理与解题表现特征。我们认为AtmosSci-Bench是推动大模型在气候服务中应用的关键一步,提供了一个标准且严谨的评估框架。源代码已开源:https://github.com/Relaxed-System-Lab/AtmosSci-Bench。
原文摘要 · Abstract (English)
The rapid advancements in large language models (LLMs), particularly in their reasoning capabilities, hold transformative potential for addressing complex challenges and boosting scientific discovery in atmospheric science. However, leveraging LLMs effectively in this domain requires a robust and comprehensive evaluation benchmark. Toward this end, we present AtmosSci-Bench, a novel benchmark designed to systematically assess LLM performance across five core categories of atmospheric science problems: hydrology, atmospheric dynamics, atmospheric physics, geophysics, and physical oceanography. AtmosSci-Bench features a dual-format design comprising both multiple-choice questions (MCQs) and open-ended questions (OEQs), enabling scalable automated evaluation alongside deeper analysis of conceptual understanding. We employ a template-based MCQ generation framework to create diverse, graduate-level problems with symbolic perturbation, while OEQs are used to probe open-ended reasoning. We conduct a comprehensive evaluation of representative LLMs, categorized into four groups: instruction-tuned models, advanced reasoning models, math-augmented models, and domain-specific climate models. Our analysis provides some interesting insights into the reasoning and problem-solving capabilities of LLMs in atmospheric science. We believe AtmosSci-Bench can serve as a critical step toward advancing LLM applications in climate services by offering a standard and rigorous evaluation framework. Our source code is available at https://github.com/Relaxed-System-Lab/AtmosSci-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。