用历史特征加速分子模拟,速度提升8倍且保持精度
BoostMD: Accelerating molecular sampling by leveraging ML force field features from previous time-steps
- 利用前步节点特征预测能量与力,简化学习任务
- 相比参考模型快8倍,支持未见二肽泛化
- 适合需要长时间、大规模分子模拟的研究者
模拟原子尺度过程(如蛋白质动力学和催化反应)对生物、化学和材料科学至关重要。机器学习力场(MLFF)已实现接近量子力学的精度并具备良好泛化能力,但其推理时间长,限制了在长时间分子动力学(MD)模拟中的应用。本文提出BoostMD,一种代理模型架构,通过复用前步计算的节点特征来预测位置变化下的能量与力,降低学习复杂度,使模型更小更快。仿真中,仅每$N$步调用高成本参考MLFF,中间步骤由轻量级BoostMD处理,计算成本可忽略。实验表明,BoostMD相比参考模型提速8倍,并能泛化至未见二肽;运行时准确采样真实玻尔兹曼分布。结合高效特征复用与精简结构,BoostMD为大规模、长时间分子模拟提供了可靠方案,使高精度ML建模更实用。
原文摘要 · Abstract (English)
Simulating atomic-scale processes, such as protein dynamics and catalytic reactions, is crucial for advancements in biology, chemistry, and materials science. Machine learning force fields (MLFFs) have emerged as powerful tools that achieve near quantum mechanical accuracy, with promising generalization capabilities. However, their practical use is often limited by long inference times compared to classical force fields, especially when running extensive molecular dynamics (MD) simulations required for many biological applications. In this study, we introduce BoostMD, a surrogate model architecture designed to accelerate MD simulations. BoostMD leverages node features computed at previous time steps to predict energies and forces based on positional changes. This approach reduces the complexity of the learning task, allowing BoostMD to be both smaller and significantly faster than conventional MLFFs. During simulations, the computationally intensive reference MLFF is evaluated only every $N$ steps, while the lightweight BoostMD model handles the intermediate steps at a fraction of the computational cost. Our experiments demonstrate that BoostMD achieves an eight-fold speedup compared to the reference model and generalizes to unseen dipeptides. Furthermore, we find that BoostMD accurately samples the ground-truth Boltzmann distribution when running molecular dynamics. By combining efficient feature reuse with a streamlined architecture, BoostMD offers a robust solution for conducting large-scale, long-timescale molecular simulations, making high-accuracy ML-driven modeling more accessible and practical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。