PRISMat高效生成晶体材料,性能超越大语言模型。
PRISMat: Policy-Driven, Permutation-Invariant Autoregressive Material Generation

- 采用置换不变架构,解决材料序列建模难题
- 推理速度更快,生成误差降低4倍
- 适合高通量材料筛选,尤其擅长表面性质预测
快速识别具备特定性能的候选材料已成为材料科学中的关键任务。机器学习作为物理模拟的替代方案,能够以更快速、低成本的方式根据稳定性和其他目标属性筛选材料,减少进入昂贵合成阶段的候选数量。近年来,大型语言模型(LLMs)被应用于该任务,但这些模型参数量大、训练和推理计算开销高,难以适用于高通量场景。这种低效源于语言模型的过度参数化以及将材料生成建模为序列学习问题的困难。本文提出PRISMat,一种成本更低、置换不变的自回归材料生成模型,有效解决了上述问题。实验表明,尽管推理时间更短,PRISMat在基于关键表面性质生成晶体薄片方面仍优于LLMs。在目标材料发现任务中,其解理能和功函数预测的平均绝对误差分别为0.188 eV/A²和2.79 eV,相较现有最佳模型误差降低4倍。
原文摘要 · Abstract (English)
Rapid identification of candidate materials with target properties has become a key task in materials science. Machine learning has emerged as an alternative to physics-based simulation, offering a faster and cheaper way to filter materials based on their stability and other target properties, reducing the number of candidates that reach the costly synthesis stage. Recently, Large Language Models (LLMs) have been applied to this role, but these models are parameter-heavy and computationally expensive both during training and at inference time, making them unsuitable for high-throughput tasks. This inefficiency stems from both the large over-parameterization of language models and the difficulty of framing material generation as a sequence learning problem. In this paper, we present PRISMat, a cost-effective, permutation-invariant model, which addresses these limitations. We show that PRISMat, despite taking less time for inference, is able to outperform LLMs in generating crystal slabs conditioned on critical materials' surface properties. In targeted material discovery, we achieve mean absolute errors of 0.188 eV/A$^2$ and 2.79 eV for cleavage energy and work function tasks, respectively, reducing the error of the next best model by 4$\times$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。