arXiv:2411.00177cond-mat.mtrl-scics.CL2024-11中稿 · NeurIPS被引 45

构建首个大规模材料性质预测LLM评测基准,推动领域标准化发展

LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction

  • 基于190万晶体结构构建多模态评测集,涵盖45种材料属性
  • 实测通用LLM在材料预测中表现不佳,需专用模型与指令微调
  • 支持零样本、少样本评估,适配研究者快速验证模型性能

大型语言模型(LLMs)在材料科学中的应用日益广泛,但针对基于LLM的材料性质预测缺乏系统性评测与标准化评估,制约了该领域的发展。本文提出LLM4Mat-Bench,目前规模最大的用于评估LLMs在晶体材料性质预测中表现的基准数据集。该基准包含约190万晶体结构,来自10个公开材料数据源,涵盖45种不同性质。其输入模态包括晶体组成、CIF文件和晶体文本描述,对应总令牌数分别为470万、6.155亿和31亿。我们利用该基准对不同规模的模型(如LLM-Prop和MatBERT)进行微调,并为类似LLM-chat的模型(如Llama、Gemma、Mistral)提供零样本和少样本提示,以评估其性质预测能力。结果表明,通用大模型在材料科学任务中面临挑战,亟需任务特定的预测模型及任务特定指令微调的LLM。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being used in materials science. However, little attention has been given to benchmarking and standardized evaluation for LLM-based materials property prediction, which hinders progress. We present LLM4Mat-Bench, the largest benchmark to date for evaluating the performance of LLMs in predicting the properties of crystalline materials. LLM4Mat-Bench contains about 1.9M crystal structures in total, collected from 10 publicly available materials data sources, and 45 distinct properties. LLM4Mat-Bench features different input modalities: crystal composition, CIF, and crystal text description, with 4.7M, 615.5M, and 3.1B tokens in total for each modality, respectively. We use LLM4Mat-Bench to fine-tune models with different sizes, including LLM-Prop and MatBERT, and provide zero-shot and few-shot prompts to evaluate the property prediction capabilities of LLM-chat-like models, including Llama, Gemma, and Mistral. The results highlight the challenges of general-purpose LLMs in materials science and the need for task-specific predictive models and task-specific instruction-tuned LLMs in materials property prediction.

材料预测大模型评测晶体结构多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。