用结构感知多模态大模型提升材料属性预测与科学推理能力
MatterChat: A Multi-Modal LLM for Material Science
- 通过桥接模块融合原子结构与文本信息,实现结构级材料数据输入
- 在材料属性预测上超越GPT-4等通用大模型,提升人机协作效率
- 适用于复杂材料合成路径设计与科学推理,推动材料研发智能化
理解与预测无机材料的性质对加速材料科学发展及在能源、电子等领域的应用至关重要。将材料结构数据与语言信息结合,通过多模态大语言模型(LLMs)有望增强人机交互。然而,如何将原子级结构以全分辨率融入LLMs仍是关键挑战。本文提出MatterChat,一种结构感知的多模态大模型,统一整合材料结构数据与文本输入。该模型采用桥接模块,有效对齐预训练的机器学习势函数与预训练大语言模型,降低训练成本并提升灵活性。实验表明,MatterChat在材料属性预测和人机交互方面显著优于GPT-4等通用大模型,并在更复杂的科学推理与分步材料合成任务中展现出实用性。
原文摘要 · Abstract (English)
Understanding and predicting the properties of inorganic materials is crucial for accelerating advancements in materials science and driving applications in energy, electronics, and beyond. Integrating material structure data with language-based information through multi-modal large language models (LLMs) offers great potential to support these efforts by enhancing human-AI interaction. However, a key challenge lies in integrating atomic structures at full resolution into LLMs. In this work, we introduce MatterChat, a versatile structure-aware multi-modal LLM that unifies material structural data and textual inputs into a single cohesive model. MatterChat employs a bridging module to effectively align a pretrained machine learning interatomic potential with a pretrained LLM, reducing training costs and enhancing flexibility. Our results demonstrate that MatterChat significantly improves performance in material property prediction and human-AI interaction, surpassing general-purpose LLMs such as GPT-4. We also demonstrate its usefulness in applications such as more advanced scientific reasoning and step-by-step material synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。