arXiv:2409.03161cs.CLcond-mat.mtrl-sci2024-09被引 1

构建材料科学领域高校水平测评数据集,评估大模型解题能力。

MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models

  • 基于大学教材构建问答对数据集,含开放回答与选择题两种形式。
  • 测试显示GPT-4在开放题中准确率显著高于GPT-3.5和ChatGPT-3.5。
  • 适合研究大模型在材料科学推理能力及提示工程影响的学者使用。

本文构建了面向大语言模型(LLMs)的材料科学领域高校水平基准数据集MaterialBENCH。该数据集包含基于大学教材的问题-答案对,分为开放回答型和多项选择型。多项选择题通过在正确答案基础上添加三个错误选项生成,使模型从中选出唯一正确答案。开放回答与选择题类型问题大部分内容重合,仅答案格式不同。我们对ChatGPT-3.5、ChatGPT-4、Bard(实验时版本)以及通过OpenAI API调用的GPT-3.5和GPT-4进行了测试。分析了各模型在不同题型下的表现差异,探讨了系统提示对多项选择题的影响。结果表明,模型在开放回答题中表现更依赖深层推理,而选择题受提示设计影响较大。该数据集有望推动大模型在复杂问题求解中的推理能力发展,助力材料研究与发现。

原文摘要 · Abstract (English)

A college-level benchmark dataset for large language models (LLMs) in the materials science field, MaterialBENCH, is constructed. This dataset consists of problem-answer pairs, based on university textbooks. There are two types of problems: one is the free-response answer type, and the other is the multiple-choice type. Multiple-choice problems are constructed by adding three incorrect answers as choices to a correct answer, so that LLMs can choose one of the four as a response. Most of the problems for free-response answer and multiple-choice types overlap except for the format of the answers. We also conduct experiments using the MaterialBENCH on LLMs, including ChatGPT-3.5, ChatGPT-4, Bard (at the time of the experiments), and GPT-3.5 and GPT-4 with the OpenAI API. The differences and similarities in the performance of LLMs measured by the MaterialBENCH are analyzed and discussed. Performance differences between the free-response type and multiple-choice type in the same models and the influence of using system massages on multiple-choice problems are also studied. We anticipate that MaterialBENCH will encourage further developments of LLMs in reasoning abilities to solve more complicated problems and eventually contribute to materials research and discovery.

材料科学大模型评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。