融合GNN与符号回归,提升材料预测准确率并实现可解释性
Combining feature-based approaches with graph neural networks and symbolic regression for synergistic performance and interpretability
- 整合多种预训练GNN的隐式特征与符号回归新特征
- 在Matbench上使MODNet误差显著降低,多任务提升超40%
- 通过代理模型将隐式特征转为物理可解释公式,适合研发人员使用
本研究提出MatterVial,一种面向材料科学的混合特征学习框架。该框架通过集成结构型(MEGNet)、组分型(ROOST)和等变型(ORB)等多种预训练图神经网络的隐式表示,结合计算高效的近似描述符及符号回归生成的新特征,拓展了特征空间。方法融合传统特征模型的化学透明性与深度学习的预测能力。在Matbench任务中,用此法增强MODNet模型,显著降低误差,性能达到甚至超越多个先进端到端GNN,在多个任务中准确率提升超过40%。集成的可解释模块利用代理模型与符号回归,将GNN提取的隐式特征还原为明确的物理意义公式。该统一框架推动材料信息学发展,提供高性能且透明的工具,符合可解释AI原则,助力更精准、自主的材料发现。
原文摘要 · Abstract (English)
This study introduces MatterVial, an innovative hybrid framework for feature-based machine learning in materials science. MatterVial expands the feature space by integrating latent representations from a diverse suite of pretrained graph neural network (GNN) models including: structure-based (MEGNet), composition-based (ROOST), and equivariant (ORB) graph networks, with computationally efficient, GNN-approximated descriptors and novel features from symbolic regression. Our approach combines the chemical transparency of traditional feature-based models with the predictive power of deep learning architectures. When augmenting the feature-based model MODNet on Matbench tasks, this method yields significant error reductions and elevates its performance to be competitive with, and in several cases superior to, state-of-the-art end-to-end GNNs, with accuracy increases exceeding 40% for multiple tasks. An integrated interpretability module, employing surrogate models and symbolic regression, decodes the latent GNN-derived descriptors into explicit, physically meaningful formulas. This unified framework advances materials informatics by providing a high-performance, transparent tool that aligns with the principles of explainable AI, paving the way for more targeted and autonomous materials discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。