arXiv:2412.10477cs.LGcond-mat.mtrl-sci2024-12被引 14

测试大模型在原子层沉积领域的科研能力,发现表现尚可但存在幻觉。

Benchmarking large language models for materials synthesis: the case of atomic layer deposition

  • 构建开放问答基准ALDbench,覆盖从研究生到专家级别的材料合成问题。
  • GPT-4o平均得分3.7(满分5),36%的问题出现低分,五处疑似幻觉。
  • 问题难度与回答质量、相关性正相关,特定性越高准确性越强。

本文提出一个开放性问题基准ALDbench,用于评估大语言模型(LLMs)在材料合成领域,特别是原子层沉积(atomic layer deposition, ALD)中的表现。该基准包含从研究生水平到领域专家级前沿问题的题目,由人类专家评审问题难度与具体性,并从整体质量、具体性、相关性和准确性四方面评估模型回答。在OpenAI GPT-4o上的测试结果显示,模型综合质量得分为3.7(1~5分制),相当于及格水平。然而,36%的问题至少有一项评分低于平均水平,深入分析发现至少五处疑似幻觉。此外,人类专家评估显示:问题难度与回答质量、相关性呈显著正相关,问题具体性与回答准确性也呈显著正相关。这表明需多维度评估LLMs,不能仅看难度或准确率。

原文摘要 · Abstract (English)

In this work we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and in particular in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI's GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1 to 5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, and the specificity of the question and the accuracy of the response as graded by the human experts. This emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

大模型评测材料合成原子层沉积幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。