测试大模型对材料知识的掌握程度,找出其适用边界。
What do Large Language Models know about materials?
- 用元素周期表测试大模型的材料知识,分析词汇与分词的影响。
- 不同开源模型在材料事实生成上表现差异显著,最高准确率仅62%。
- 为材料工程应用提供选择依据,明确哪些环节需专用模型。
大型语言模型(LLMs)在机械工程和材料科学领域应用日益广泛。作为通过语言建立关联的模型,它们可沿材料科学中的加工-结构-性能-性能链进行逐步推理。当前的LLMs主要基于互联网数据训练,但互联网内容以非科学信息为主。若将LLMs用于工程场景,有必要考察其内在知识——即生成材料正确信息的能力。本文以元素周期表为例,揭示词汇与分词对材料指纹独特性的作用,并评估多个先进开源模型生成材料事实的准确性。结果表明,不同模型表现差异明显,最高准确率为62%。研究构建了材料知识基准,明确了在PSPP链中哪些步骤可使用通用模型,哪些环节需依赖专用模型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly applied in the fields of mechanical engineering and materials science. As models that establish connections through the interface of language, LLMs can be applied for step-wise reasoning through the Processing-Structure-Property-Performance chain of material science and engineering. Current LLMs are built for adequately representing a dataset, which is the most part of the accessible internet. However, the internet mostly contains non-scientific content. If LLMs should be applied for engineering purposes, it is valuable to investigate models for their intrinsic knowledge -- here: the capacity to generate correct information about materials. In the current work, for the example of the Periodic Table of Elements, we highlight the role of vocabulary and tokenization for the uniqueness of material fingerprints, and the LLMs' capabilities of generating factually correct output of different state-of-the-art open models. This leads to a material knowledge benchmark for an informed choice, for which steps in the PSPP chain LLMs are applicable, and where specialized models are required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。