首个生态环保领域大模型评测数据集,助力AI在环境应用中精准评估。
Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment Domain
- 构建16个环境主题的1130组问答对,覆盖不同难度与类型。
- 提供统一评估框架,实现大模型在环境任务中的客观对比。
- 适合研究可持续发展、环境AI评估的学者和开发者使用。
生成式AI在生态监测、数据分析、教育和政策支持等领域具有巨大潜力,但其效果受限于缺乏统一的评估框架。为此,我们提出了首个面向生态与环境科学的大语言模型评估基准——环境大模型评测数据集(ELLE)。该数据集包含1,130组问答对,覆盖16个环境主题,按领域、难度和类型进行分类。通过标准化评估流程,ELLE实现了对生成式AI在环境领域性能的一致性与客观性比较,推动了可持续环境应用中AI技术的发展。数据集与代码已开源,可访问 https://elle.ceeai.net/ 及 https://github.com/CEEAI/elle。
原文摘要 · Abstract (English)
Generative AI holds significant potential for ecological and environmental applications such as monitoring, data analysis, education, and policy support. However, its effectiveness is limited by the lack of a unified evaluation framework. To address this, we present the Environmental Large Language model Evaluation (ELLE) question answer (QA) dataset, the first benchmark designed to assess large language models and their applications in ecological and environmental sciences. The ELLE dataset includes 1,130 question answer pairs across 16 environmental topics, categorized by domain, difficulty, and type. This comprehensive dataset standardizes performance assessments in these fields, enabling consistent and objective comparisons of generative AI performance. By providing a dedicated evaluation tool, ELLE dataset promotes the development and application of generative AI technologies for sustainable environmental outcomes. The dataset and code are available at https://elle.ceeai.net/ and https://github.com/CEEAI/elle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。