arXiv:2505.18653cs.CL2025-05中稿 · ACL被引 3

构建首个覆盖25项气候任务的NLP评测基准,评估大模型在气候变化领域的表现。

Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change

  • 整合13个数据集,新增新闻分类数据集,覆盖文本分类等关键任务
  • 评估2B至70B参数开源大模型在零样本与少样本下的性能表现
  • 为气候相关NLP研究提供标准化评测工具,适合关注环境智能的研究者

Climate-Eval是一个全面的基准测试,用于评估自然语言处理模型在气候变化相关任务中的表现。该基准整合了现有数据集,并新增一个专为本次发布设计的新闻分类数据集,共涵盖25项任务,基于13个数据集,覆盖文本分类、问答和信息抽取等气候话语核心领域。本基准为系统评估大语言模型(LLMs)在这些任务上的表现提供了标准化评估套件。此外,我们对开源大模型(参数量从2B到70B)进行了广泛评估,涵盖零样本和少样本设置,分析其在气候变化领域的优势与局限性。

原文摘要 · Abstract (English)

Climate-Eval is a comprehensive benchmark designed to evaluate natural language processing models across a broad range of tasks related to climate change. Climate-Eval aggregates existing datasets along with a newly developed news classification dataset, created specifically for this release. This results in a benchmark of 25 tasks based on 13 datasets, covering key aspects of climate discourse, including text classification, question answering, and information extraction. Our benchmark provides a standardized evaluation suite for systematically assessing the performance of large language models (LLMs) on these tasks. Additionally, we conduct an extensive evaluation of open-source LLMs (ranging from 2B to 70B parameters) in both zero-shot and few-shot settings, analyzing their strengths and limitations in the domain of climate change.

气候AI评测基准大模型NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。