提出高效并行算法,让分子动力学中MPNN模型突破百亿原子规模限制。
Efficient Parallelization of Message Passing Neural Network Potentials for Large-scale Molecular Dynamics
- 通过减少层间通信开销,实现MPNN模型跨节点线性扩展。
- 在百万级原子系统上达到与局部模型相当的计算速度。
- 通用框架可适配多种MPNN模型,适合大规模复杂体系模拟。
机器学习势能在加速原子模拟方面取得显著进展。许多基于原子中心局部描述符的方法天然适合并行计算。近年来,消息传递神经网络(MPNN)模型因其卓越精度而日益流行。然而,如何在多节点间高效并行化MPNN仍具挑战,限制了其在大规模模拟中的实际应用。本文提出一种高效的MPNN并行算法,在每层消息传递中仅对局部原子进行最小化数据通信,避免冗余计算,使计算开销随层数线性增长。结合递归嵌入原子神经网络模型,该算法在多个基准系统中展现出优异的强可扩展性和弱可扩展性。该方法使基于MPNN的分子动力学模拟在超过1亿原子规模下,运行速度接近纯局部模型,极大拓展了MPNN势函数的应用边界。该通用并行框架可赋能各类MPNN模型,高效模拟超大且复杂的系统。
原文摘要 · Abstract (English)
Machine learning potentials have achieved great success in accelerating atomistic simulations. Many of them relying on atom-centered local descriptors are natural for parallelization. More recent message passing neural network (MPNN) models have demonstrated their superior accuracy and become increasingly popular. However, efficiently parallelizing MPNN models across multiple nodes remains challenging, limiting their practical applications in large-scale simulations. Here, we propose an efficient parallel algorithm for MPNN models, in which additional data communication is minimized among local atoms only in each MP layer without redundant computation, thus scaling linearly with the layer number. Integrated with our recursively embedded atom neural network model, this algorithm demonstrates excellent strong scaling and weak scaling behaviors in several benchmark systems. This approach enables massive molecular dynamics simulations on MPNN models as fast as on strictly local models for over 100 million atoms, vastly extending the applicability of the MPNN potential to an unprecedented scale. This general parallelization framework can empower various MPNN models to efficiently simulate very large and complex systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。