用经典力场数据预训练,显著降低量子计算数据需求。
Autotuning T-PaiNN: Enabling Data-Efficient GNN Interatomic Potential Development via Classical-to-Quantum Transfer Learning
- 先用经典力场数据预训练,再用少量量子数据微调。
- 低数据下误差降低25倍,收敛速度更快。
- 适合缺乏量子数据的研究者快速构建高精度模型。
机器学习势能(MLIP)特别是基于图神经网络(GNN)的模型,可在大幅降低计算成本的同时实现接近密度泛函理论(DFT)的精度。然而其实际应用常受限于昂贵的量子力学训练数据需求。本文提出迁移学习框架T-PaiNN,通过利用廉价的经典力场数据提升GNN-MLIP的数据效率。方法包括:在大规模经典分子模拟数据上预训练PaiNN模型,再用少量DFT数据进行微调(称作自动调优)。在气相分子系统(QM9数据集)和凝聚相液态水体系中均验证了有效性。所有情况下,T-PaiNN均显著优于仅使用DFT数据训练的模型,在低数据条件下平均绝对误差降低达一个数量级,且训练收敛加速。例如在QM9上低数据下误差减少最高达25倍;液态水模拟中能量、力及密度、扩散等实验相关性质预测均获提升。性能提升源于模型从大量经典采样中学习到势能面的通用特征,并进一步精炼至量子精度。本工作确立了从经典力场迁移学习为开发高精度、数据高效GNN势能的有效策略,推动MLIP在复杂化学体系中的广泛应用。
原文摘要 · Abstract (English)
Machine-learned interatomic potentials (MLIPs), particularly graph neural network (GNN)-based models, offer a promising route to achieving near-density functional theory (DFT) accuracy at significantly reduced computational cost. However, their practical deployment is often limited by the large volumes of expensive quantum mechanical training data required. In this work, we introduce a transfer learning framework, Transfer-PaiNN (T-PaiNN), that substantially improves the data efficiency of GNN-MLIPs by leveraging inexpensive classical force field data. The approach consists of pretraining a PaiNN MLIP architecture on large-scale datasets generated from classical molecular simulations, followed by fine-tuning (dubbed autotuning) using a comparatively small DFT dataset. We demonstrate the effectiveness of autotuning T-PaiNN on both gas-phase molecular systems (QM9 dataset) and condensed-phase liquid water. Across all cases, T-PaiNN significantly outperforms models trained solely on DFT data, achieving order-of-magnitude reductions in mean absolute error while accelerating training convergence. For example, using the QM9 data set, error reductions of up to 25 times are observed in low-data regimes, while liquid water simulations show improved predictions of energies, forces, and experimentally relevant properties such as density and diffusion. These gains arise from the model's ability to learn general features of the potential energy surface from extensive classical sampling, which are subsequently refined to quantum accuracy. Overall, this work establishes transfer learning from classical force fields as a practical and computationally efficient strategy for developing high-accuracy, data-efficient GNN interatomic potentials, enabling broader application of MLIPs to complex chemical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。