构建首个覆盖地球科学全领域的LLM评估基准,测试模型从基础到探索性研究的能力。
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
- 基于十万篇论文构建三类问答数据集,覆盖五大地球圈层与114个学科
- 引入开放对话数据集Earth-Gold,评估模型提出方法、分析局限等高级科研能力
- 发现11个主流大模型在科学探索中普遍存在短板,适合评估AI科研潜力
大型语言模型(LLMs)在科学领域的应用日益广泛,亟需专门的评估基准。现有基准或泛化科学知识缺乏地球科学针对性,或仅覆盖单一子领域,且普遍忽略对开放科学探索能力的评测。本文提出一个全面专业的地球科学评估基准EarthSE,涵盖从基础到高级的科学探索能力。基于10万篇研究论文,构建两个问答数据集:Earth-Iron(广度覆盖)和Earth-Silver(高难度专业深度)。二者覆盖五大地球圈层、114个学科和11类任务,评估基础科学知识。最核心的是引入Earth-Gold数据集,包含开放式的多轮对话,用于评测模型的方法推导、局限分析与概念创新等高级科研能力。实验表明,11个主流大模型在不同任务中均存在明显不足,显示其科学探索能力仍有巨大提升空间。基准已开源:https://huggingface.co/ai-earth。
原文摘要 · Abstract (English)
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated subdomains, lacking holistic evaluation. Furthermore, current benchmarks typically neglect the assessment of LLMs' capabilities in open-ended scientific exploration. In this paper, we present a comprehensive and professional benchmark for the Earth sciences, designed to evaluate the capabilities of LLMs in scientific exploration within this domain, spanning from fundamental to advanced levels. Leveraging a corpus of 100,000 research papers, we first construct two Question Answering (QA) datasets: Earth-Iron, which offers extensive question coverage for broad assessment, and Earth-Silver, which features a higher level of difficulty to evaluate professional depth. These datasets encompass five Earth spheres, 114 disciplines, and 11 task categories, assessing foundational knowledge crucial for scientific exploration. Most notably, we introduce Earth-Gold with new metrics, a dataset comprising open-ended multi-turn dialogues specifically designed to evaluate the advanced capabilities of LLMs in scientific exploration, including methodology induction, limitation analysis, and concept proposal. Extensive experiments reveal limitations in 11 leading LLMs across different domains and tasks, highlighting considerable room for improvement in their scientific exploration capabilities. The benchmark is available on https://huggingface.co/ai-earth .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。