用注意力机制提升神经网络势函数的可扩展性,实现更快更准的跨化学体系模拟。
The Importance of Being Scalable: Improving the Speed and Accuracy of Neural Network Interatomic Potentials Across Chemical Domains
- 通过多头自注意力机制增强图神经网络表达能力,突破传统物理约束对模型扩展的限制。
- 相比现有方法,推理速度提升至少10倍,内存占用减少5倍,性能达多个数据集新高。
- 为通用神经网络势函数设计提供新范式,适合追求高效与泛化的材料模拟研究者。
规模扩展在机器学习中对提升模型性能和泛化能力至关重要,但其在神经网络原子间势(NNIP)领域的研究仍不充分。现有主流方法依赖旋转等变性等复杂物理约束,限制了模型的可扩展性并可能导致性能瓶颈。本文系统研究了NNIP的缩放策略,发现基于注意力机制的扩展方式既高效又能增强表达能力。据此提出可高效扩展的原子间势模型EScAIP:在图神经网络中引入邻域级多头自注意力,并配合高度优化的GPU注意力内核。实验表明,相较于现有模型,EScaIP推理速度提升至少10倍,内存占用减少5倍,在催化剂(OC20、OC22)、分子(SPICE)及材料(MPTrj)等多个数据集上达到当前最优表现。该方法代表一种新范式,强调通过规模化实现更高表达力,可随计算资源和训练数据持续扩展。
原文摘要 · Abstract (English)
Scaling has been critical in improving model performance and generalization in machine learning. It involves how a model's performance changes with increases in model size or input data, as well as how efficiently computational resources are utilized to support this growth. Despite successes in other areas, the study of scaling in Neural Network Interatomic Potentials (NNIPs) remains limited. NNIPs act as surrogate models for ab initio quantum mechanical calculations. The dominant paradigm here is to incorporate many physical domain constraints into the model, such as rotational equivariance. We contend that these complex constraints inhibit the scaling ability of NNIPs, and are likely to lead to performance plateaus in the long run. In this work, we take an alternative approach and start by systematically studying NNIP scaling strategies. Our findings indicate that scaling the model through attention mechanisms is efficient and improves model expressivity. These insights motivate us to develop an NNIP architecture designed for scalability: the Efficiently Scaled Attention Interatomic Potential (EScAIP). EScAIP leverages a multi-head self-attention formulation within graph neural networks, applying attention at the neighbor-level representations. Implemented with highly-optimized attention GPU kernels, EScAIP achieves substantial gains in efficiency--at least 10x faster inference, 5x less memory usage--compared to existing NNIPs. EScAIP also achieves state-of-the-art performance on a wide range of datasets including catalysts (OC20 and OC22), molecules (SPICE), and materials (MPTrj). We emphasize that our approach should be thought of as a philosophy rather than a specific model, representing a proof-of-concept for developing general-purpose NNIPs that achieve better expressivity through scaling, and continue to scale efficiently with increased computational resources and training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。