评测大模型在材料科学工具上的代码生成能力,发现通用模型更优且越简单越好。
MatTools: Benchmarking Large Language Models for Materials Science Tools
- 构建双轨基准:问答对与真实代码生成任务
- 覆盖69,225组问答和138个子任务,测试代码生成准确性
- 揭示通用模型优于专用模型,简单方法更有效
大型语言模型(LLMs)在材料科学领域应用日益广泛,涵盖文献理解、性质预测、材料发现和合金设计。与此同时,多种基于物理的计算方法被用于材料性质计算。本文提出MatTools基准,评估LLMs通过生成并安全执行物理基础材料科学软件代码来回答材料科学问题的能力。该框架包含两个互补组件:基于pymatgen代码库与文档的材料模拟工具问答基准(含69,225组问答对),以及包含49个任务(138个子任务)的真实世界工具使用基准,要求生成可运行的Python代码进行材料性质计算。对多种LLM的评估得出三个关键结论:(1)通用模型表现优于专用模型;(2)AI了解自身局限性;(3)简单方法更有效。MatTools为材料科学中LLM工具应用提供标准化评估与改进框架,推动更高效的材料与通用科研人工智能系统发展。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied to materials science questions, including literature comprehension, property prediction, materials discovery and alloy design. At the same time, a wide range of physics-based computational approaches have been developed in which materials properties can be calculated. Here, we propose a benchmark application to evaluate the proficiency of LLMs to answer materials science questions through the generation and safe execution of codes based on such physics-based computational materials science packages. MatTools is built on two complementary components: a materials simulation tool question-answer (QA) benchmark and a real-world tool-usage benchmark. We designed an automated methodology to efficiently collect real-world materials science tool-use examples. The QA benchmark, derived from the pymatgen (Python Materials Genomics) codebase and documentation, comprises 69,225 QA pairs that assess the ability of an LLM to understand materials science tools. The real-world benchmark contains 49 tasks (138 subtasks) requiring the generation of functional Python code for materials property calculations. Our evaluation of diverse LLMs yields three key insights: (1)Generalists outshine specialists;(2)AI knows AI; and (3)Simpler is better. MatTools provides a standardized framework for assessing and improving LLM capabilities for materials science tool applications, facilitating the development of more effective AI systems for materials science and general scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。