用现成模型做老师,加速高精度分子模拟,省下10倍算力。
Knowledge Distillation Framework for Accelerating High-Accuracy Neural Network-Based Molecular Dynamics Simulations
- 不用微调教师模型,用现成预训练模型生成数据
- 减少10倍密度泛函理论计算,保持精度不变
- 学生模型推理快106倍,适合大规模分子动力学
神经网络势函数(NNPs)为分子动力学(MD)模拟提供了强大替代方案。准确稳定的模拟需涵盖低能稳定结构和高能结构的训练数据。传统知识蒸馏方法通过微调教师模型生成学生模型训练数据,但在材料类模型中会增加能量壁垒,难以获取高能结构。为此,我们提出一种新框架,采用未经微调的现成预训练模型作为教师,其更平缓的能量景观有助于探索更广结构范围,包括关键的高能结构。该框架分两阶段训练:首先用教师生成的数据训练学生模型;再用少量高精度密度泛函理论(DFT)数据进行微调。在聚乙二醇(有机)和L₁₀GeP₂S₁₂(无机)材料上验证,性能优于或相当现有方法。重要的是,相比现有方法,本方法将昂贵的DFT计算减少10倍,且学生模型推理速度提升至教师模型的106倍,显著加快模拟效率。
原文摘要 · Abstract (English)
Neural network potentials (NNPs) offer a powerful alternative to traditional force fields for molecular dynamics (MD) simulations. Accurate and stable MD simulations, crucial for evaluating material properties, require training data encompassing both low-energy stable structures and high-energy structures. Conventional knowledge distillation (KD) methods fine-tune a pre-trained NNP as a teacher model to generate training data for a student model. However, in material-specific models, this fine-tuning process increases energy barriers, making it difficult to create training data containing high-energy structures. To address this, we propose a novel KD framework that leverages a non-fine-tuned, off-the-shelf pre-trained NNP as a teacher. Its gentler energy landscape facilitates the exploration of a wider range of structures, including the high-energy structures crucial for stable MD simulations. Our framework employs a two-stage training process: first, the student NNP is trained with a dataset generated by the off-the-shelf teacher; then, it is fine-tuned with a smaller, high-accuracy density functional theory (DFT) dataset. We demonstrate the effectiveness of our framework by applying it to both organic (polyethylene glycol) and inorganic (L$_{10}$GeP$_{2}$S$_{12}$) materials, achieving comparable or superior accuracy in reproducing physical properties compared to existing methods. Importantly, our method reduces the number of expensive DFT calculations by 10x compared to existing NNP generation methods, without sacrificing accuracy. Furthermore, the resulting student NNP achieves up to 106x speedup in inference compared to the teacher NNP, enabling significantly faster and more efficient MD simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。