arXiv:2412.11970cs.CL2024-12被引 14

用大模型直接读论文和数据,自动预测材料性能

DARWIN 1.5: Large Language Models as Materials Science Adapted Learners

  • 用自然语言输入替代复杂描述符,统一处理不同材料任务
  • 在8个材料设计任务中准确率比基础模型提升59.1%
  • 适合材料科研人员快速探索新配方和性质

材料发现与设计旨在复杂多样的物理空间中寻找具有理想特性的成分与结构。传统方法如高通量模拟或机器学习常依赖复杂描述符,限制了跨体系的泛化与迁移能力,且难以表征受结构缺陷和成分变化影响的宏观材料性能,降低实际应用价值。为此,我们提出DARWIN 1.5,目前最大开源面向材料科学的大型语言模型。通过使用自然语言作为输入,DARWIN无需任务特定描述符,实现材料性质预测与发现的灵活统一方法。模型融合600万篇材料领域论文及来自49,256种材料的21个实验数据集,涵盖多种模态,并支持跨任务知识迁移。增强版模型在预测准确率上相较基线LLaMA-7B架构最高提升59.1%,并在8项材料设计任务中超越现有最优机器学习方法。结果表明,大语言模型可成为材料科学中通用、可扩展模型的有力基础。

原文摘要 · Abstract (English)

Materials discovery and design aim to find compositions and structures with desirable properties over highly complex and diverse physical spaces. Traditional solutions, such as high-throughput simulations or machine learning, often rely on complex descriptors, which hinder generalizability and transferability across different material systems. Moreover, These descriptors may inadequately represent macro-scale material properties, which are influenced by structural imperfections and compositional variations in real-world samples, thus limiting their practical applicability. To address these challenges, we propose DARWIN 1.5, the largest open-source large language model tailored for materials science. By leveraging natural language as input, DARWIN eliminates the need for task-specific descriptors and enables a flexible, unified approach to material property prediction and discovery. Our approach integrates 6M material domain papers and 21 experimental datasets from 49,256 materials across modalities while enabling cross-task knowledge transfer. The enhanced model achieves up to 59.1% improvement in prediction accuracy over the base LLaMA-7B architecture and outperforms SOTA machine learning approaches across 8 materials design tasks. These results establish LLMs as a promising foundation for developing versatile and scalable models in materials science.

材料科学大模型属性预测生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。