提升深度等变势能模型训练与推理速度,支持大规模分子模拟。
High-performance training and inference for deep equivariant interatomic potentials
- 重构NequIP框架,支持多节点并行与PyTorch 2.0编译加速。
- 在SPICE 2数据集上训练Allegro模型,推理速度提升最高18倍。
- 首次实现端到端PyTorch AOT编译,适合高性能分子动力学研究者。
基于深度等变神经网络的机器学习原子间势能,在分子动力学和高通量筛选等原子建模任务中表现出顶尖的精度与计算效率。随着数据集规模和下游工作流需求迅速增长,稳健可扩展的软件成为关键。本文对NequIP框架进行重大重构,聚焦多节点并行、计算性能与可扩展性。新框架支持大规模数据集分布式训练,并突破了训练时完全利用PyTorch 2.0编译器的障碍。通过案例研究,在有机分子系统SPICE 2数据集上训练Allegro模型,验证了显著加速效果。针对推理,首次构建端到端使用PyTorch Ahead-of-Time Inductor编译器的基础设施。此外,为Allegro模型中最耗时的张量积操作实现定制内核。上述改进使实际系统规模下的分子动力学计算速度提升最高达18倍。
原文摘要 · Abstract (English)
Machine learning interatomic potentials, particularly those based on deep equivariant neural networks, have demonstrated state-of-the-art accuracy and computational efficiency in atomistic modeling tasks like molecular dynamics and high-throughput screening. The size of datasets and demands of downstream workflows are growing rapidly, making robust and scalable software essential. This work presents a major overhaul of the NequIP framework focusing on multi-node parallelism, computational performance, and extensibility. The redesigned framework supports distributed training on large datasets and removes barriers preventing full utilization of the PyTorch 2.0 compiler at train time. We demonstrate this acceleration in a case study by training Allegro models on the SPICE 2 dataset of organic molecular systems. For inference, we introduce the first end-to-end infrastructure that uses the PyTorch Ahead-of-Time Inductor compiler for machine learning interatomic potentials. Additionally, we implement a custom kernel for the Allegro model's most expensive operation, the tensor product. Together, these advancements speed up molecular dynamics calculations on system sizes of practical relevance by up to a factor of 18.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。