ChemEval为化学领域大模型提供多层级评测,精准衡量其真实科研能力。
ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
- 构建4个化学层级、12维评估体系,涵盖42项专家设计任务
- 通用模型在文献理解上强,但复杂化学任务表现不足
- 专用模型化学能力更强,适合科研人员验证模型实际价值
近年来,大语言模型在化学领域的应用日益广泛,推动了面向化学任务的基准测试发展。然而,现有基准难以满足化学研究者的真实需求。为此,本文提出 extbf{ extit{ChemEval}},对大语言模型在化学领域的综合能力进行系统评估。该评估体系基于开放数据与化学专家精心构建的数据,识别出化学中的4个关键递进层级,覆盖12个维度和42项不同类型的化学任务,确保任务具有实际科研价值。实验中,在零样本和少样本学习场景下,对12个主流大模型进行了评测,采用精心挑选的示范样例与提示设计。结果表明,通用模型如GPT-4和Claude-3.5在文献理解与指令遵循方面表现优异,但在需要高级化学知识的任务中表现不佳;而专用模型虽在化学能力上有所提升,但文学理解能力下降。这说明大模型在应对复杂化学任务时仍有巨大改进空间。本工作将为推动化学领域大模型发展提供重要参考。相关基准与分析将在{ extcolor{blue}{ exttt{https://github.com/USTC-StarTeam/ChemEval}}}公开。
原文摘要 · Abstract (English)
There is a growing interest in the role that LLMs play in chemistry which lead to an increased focus on the development of LLMs benchmarks tailored to chemical domains to assess the performance of LLMs across a spectrum of chemical tasks varying in type and complexity. However, existing benchmarks in this domain fail to adequately meet the specific requirements of chemical research professionals. To this end, we propose \textbf{\textit{ChemEval}}, which provides a comprehensive assessment of the capabilities of LLMs across a wide range of chemical domain tasks. Specifically, ChemEval identified 4 crucial progressive levels in chemistry, assessing 12 dimensions of LLMs across 42 distinct chemical tasks which are informed by open-source data and the data meticulously crafted by chemical experts, ensuring that the tasks have practical value and can effectively evaluate the capabilities of LLMs. In the experiment, we evaluate 12 mainstream LLMs on ChemEval under zero-shot and few-shot learning contexts, which included carefully selected demonstration examples and carefully designed prompts. The results show that while general LLMs like GPT-4 and Claude-3.5 excel in literature understanding and instruction following, they fall short in tasks demanding advanced chemical knowledge. Conversely, specialized LLMs exhibit enhanced chemical competencies, albeit with reduced literary comprehension. This suggests that LLMs have significant potential for enhancement when tackling sophisticated tasks in the field of chemistry. We believe our work will facilitate the exploration of their potential to drive progress in chemistry. Our benchmark and analysis will be available at {\color{blue} \url{https://github.com/USTC-StarTeam/ChemEval}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。