用廉价量子计算替代高成本数据,实现高精度分子动力学模拟
Implicit Delta Learning of High Fidelity Neural Network Potentials
- 通过多任务架构融合不同精度量子计算数据
- 仅需1/50的高精度数据即达同等精度
- 适合材料与药物研发中的大规模分子模拟
神经网络势能(NNPs)为分子动力学(MD)模拟提供了快速且精确的替代方案,但其训练依赖于高精度量子力学(QM)方法生成的昂贵数据。本文提出隐式增量学习(IDLe)方法,通过利用更廉价的半经验量子计算数据,显著降低对高精度QM数据的需求,同时保持NNP的精度和推理效率。IDLe采用端到端多任务架构,基于共享的原子系统潜在表示,通过特定保真度的输出头解码能量。在多种场景下,IDLe在仅使用高精度数据的1/50时,仍能达到单一致高保真基线的精度。该方法可大幅降低数据生成成本,提升模型准确率与泛化能力,拓展化学覆盖范围,推动材料科学与药物发现中的分子动力学研究。此外,我们还提供了一个包含1100万条半经验量子计算结果的新数据集,以支持未来的多保真度NNP建模。
原文摘要 · Abstract (English)
Neural network potentials (NNPs) offer a fast and accurate alternative to ab-initio methods for molecular dynamics (MD) simulations but are hindered by the high cost of training data from high-fidelity Quantum Mechanics (QM) methods. Our work introduces the Implicit Delta Learning (IDLe) method, which reduces the need for high-fidelity QM data by leveraging cheaper semi-empirical QM computations without compromising NNP accuracy or inference cost. IDLe employs an end-to-end multi-task architecture with fidelity-specific heads that decode energies based on a shared latent representation of the input atomistic system. In various settings, IDLe achieves the same accuracy as single high-fidelity baselines while using up to 50x less high-fidelity data. This result could significantly reduce data generation cost and consequently enhance accuracy and generalization, and expand chemical coverage for NNPs, advancing MD simulations for material science and drug discovery. Additionally, we provide a novel set of 11 million semi-empirical QM calculations to support future multi-fidelity NNP modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。