用主动学习动态补全分子模拟中的未知构象,提升粗粒度模型精度。
Active Learning for Machine Learning Driven Molecular Dynamics
- 基于RMSD筛选模拟帧,主动查询原子级数据填补模型盲区。
- 在自研基准上使W1指标提升33.05%,有效覆盖新构象空间。
- 适合需要长期稳定模拟的生物分子研究者使用。
机器学习驱动的粗粒度(CG)势能模型虽快速,但在模拟触及未采样生物分子构象时会性能退化;而生成广泛原子级(AA)数据以应对这一问题计算成本过高。本文提出一种新型主动学习(AL)框架,用于分子动力学(MD)中CG神经网络势能的训练。基于CGSchNet模型,该方法通过分析MD模拟中帧的均方根偏差(RMSD),实时识别需补充数据的区域,并在训练过程中向“专家”查询原子级数据。此框架在保持粗粒度效率的同时,精准修正模型在构象空间中的覆盖率缺口。实验表明,基于该框架训练的CGSchNet模型在自研基准套件上对Chignolin蛋白进行模拟时,时间延迟独立分量分析(TICA)空间的Wasserstein-1(W1)指标提升33.05%。
原文摘要 · Abstract (English)
Machine-learned coarse-grained (CG) potentials are fast, but degrade over time when simulations reach under-sampled bio-molecular conformations, and generating widespread all-atom (AA) data to combat this is computationally infeasible. We propose a novel active learning (AL) framework for CG neural network potentials in molecular dynamics (MD). Building on the CGSchNet model, our method employs root mean squared deviation (RMSD)-based frame selection from MD simulations in order to generate data on-the-fly by querying an oracle during the training of a neural network potential. This framework preserves CG-level efficiency while correcting the model at precise, RMSD-identified coverage gaps. By training CGSchNet, a coarse-grained neural network potential, we empirically show that our framework explores previously unseen configurations and trains the model on unexplored regions of conformational space. Our active learning framework enables a CGSchNet model trained on the Chignolin protein to achieve a 33.05\% improvement in the Wasserstein-1 (W1) metric in Time-lagged Independent Component Analysis (TICA) space on an in-house benchmark suite.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。