突破大分子电子结构预测的计算瓶颈,实现超大规模系统高效模拟。
Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction
- 通过图分区与直接GPU通信,降低跨卡数据交换开销。
- 在512张GPU上实现87%并行效率,支持3000至19万原子体系。
- 适用于含缺陷、界面或无序相材料的电子性质研究。
基于密度泛函理论(DFT)数据训练的等变图神经网络(eGNN)有望在前所未有的规模下实现电子结构预测,从而研究具有扩展缺陷、界面或无序相材料的电子特性。然而,由于原子轨道相互作用通常延伸超过10埃,所需图表示往往高度连通,导致训练和推理的内存需求超出现代GPU容量。本文提出一种分布式eGNN实现,利用直接GPU通信,并引入输入图的划分策略,减少跨GPU的嵌入交换次数。该方案在Alps超算上展现出强可扩展性(最高128张GPU),弱可扩展性达512张GPU,对3,000至190,000原子体系保持87%的并行效率。
原文摘要 · Abstract (English)
Equivariant Graph Neural Networks (eGNNs) trained on density-functional theory (DFT) data can potentially perform electronic structure prediction at unprecedented scales, enabling investigation of the electronic properties of materials with extended defects, interfaces, or exhibiting disordered phases. However, as interactions between atomic orbitals typically extend over 10+ angstroms, the graph representations required for this task tend to be densely connected, and the memory requirements to perform training and inference on these large structures can exceed the limits of modern GPUs. Here we present a distributed eGNN implementation which leverages direct GPU communication and introduce a partitioning strategy of the input graph to reduce the number of embedding exchanges between GPUs. Our implementation shows strong scaling up to 128 GPUs, and weak scaling up to 512 GPUs with 87% parallel efficiency for structures with 3,000 to 190,000 atoms on the Alps supercomputer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。