构建气候问答评估框架,提升大模型在气候科学中的可信度。
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
- 基于教材与专家协作生成问答对,确保内容科学性。
- 推出专家标注的黄金数据集与大规模合成数据集。
- 提供评估策略,适合研究气候大模型可靠性的团队使用。
大型语言模型(LLMs)在气候科学领域的应用日益受到关注,但缺乏全面的评估框架来衡量其输出质量与科学合理性。为此,我们提出ClimaGen(气候问答生成器),一种结合研究生教材与气候科学家协作的自适应学习框架,生成高质量问答对。由此构建了ClimaQA-Gold——一个专家标注的基准数据集,以及ClimaQA-Silver——一个大规模、综合性合成问答数据集。最后,我们设计评估策略,并在多个大型语言模型上进行测试。结果揭示了增强气候领域大模型知识的不同方法的新见解。代码已开源,地址为 https://github.com/Rose-STL-Lab/genie-climaqa。
原文摘要 · Abstract (English)
The use of Large Language Models (LLMs) in climate science has recently gained significant attention. However, a critical issue remains: the lack of a comprehensive evaluation framework capable of assessing the quality and scientific validity of model outputs. To address this issue, we develop ClimaGen (Climate QA Generator), an adaptive learning framework that generates question-answer pairs from graduate textbooks with climate scientists in the loop. As a result, we present ClimaQA-Gold, an expert-annotated benchmark dataset alongside ClimaQA-Silver, a large-scale, comprehensive synthetic QA dataset for climate science. Finally, we develop evaluation strategies and compare different LLMs on our benchmarks. Our results offer novel insights into various approaches used to enhance knowledge of climate LLMs. The source code is publicly available at https://github.com/Rose-STL-Lab/genie-climaqa
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。