用图文交叉注意力提升材料性质预测准确率
CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction
- 通过交叉注意力融合晶体图结构与文本描述的细粒度特征
- 在四种材料属性上平均降低10.2%~35.7%的预测误差
- 适合关注材料科学多模态建模的研究者
图神经网络(GNN)在将晶体结构建模为图以预测材料性质方面取得显著进展,但通常难以捕捉全局结构特征(如晶系),限制了预测性能。为此,我们提出CAST——一种基于交叉注意力的多模态模型,将图表示与材料文本描述相结合,有效保留关键结构与成分信息。不同于依赖材料级聚合嵌入的CrysMMNet和MultiMat等方法,CAST利用交叉注意力机制融合节点级图特征与文本词元级特征。此外,引入掩码节点预测预训练策略,进一步增强节点与文本嵌入的对齐。实验表明,CAST在四个关键材料性质(形成能、带隙、体弹模量、剪切模量)上均优于现有基线模型,平均相对MAE提升10.2%至35.7%。注意力图分析证实预训练对多模态表示对齐的重要性。本研究凸显了多模态学习框架在构建更精准、全局感知的材料预测模型中的潜力。
原文摘要 · Abstract (English)
Recent advancements in graph neural networks (GNNs) have significantly enhanced the prediction of material properties by modeling crystal structures as graphs. However, GNNs often struggle to capture global structural characteristics, such as crystal systems, limiting their predictive performance. To overcome this issue, we propose CAST, a cross-attention-based multimodal model that integrates graph representations with textual descriptions of materials, effectively preserving critical structural and compositional information. Unlike previous approaches, such as CrysMMNet and MultiMat, which rely on aggregated material-level embeddings, CAST leverages cross-attention mechanisms to combine fine-grained graph node-level and text token-level features. Additionally, we introduce a masked node prediction pretraining strategy that further enhances the alignment between node and text embeddings. Our experimental results demonstrate that CAST outperforms existing baseline models across four key material properties-formation energy, band gap, bulk modulus, and shear modulus-with average relative MAE improvements ranging from 10.2% to 35.7%. Analysis of attention maps confirms the importance of pretraining in effectively aligning multimodal representations. This study underscores the potential of multimodal learning frameworks for developing more accurate and globally informed predictive models in materials science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。