用能量模型修正分子动力学采样不一致问题
Consistent Sampling and Simulation: Molecular Dynamics with Energy-Based Diffusion Models
- 基于福克-普朗克方程设计正则化项,确保得分模型与能量分布一致
- 在多肽和快速折叠蛋白系统上实现高效采样与模拟,性能优于现有方法
- 提供可迁移的玻尔兹曼模拟器,适合生物分子动力学研究者使用
近年来,基于平衡分子分布训练的扩散模型在生物分子采样中表现良好。除了直接采样外,该模型的得分还可用于推导作用于分子系统的力。然而,经典扩散采样虽能恢复训练分布,其对应的能量解释常与分布不一致,即使在低维模型系统中也是如此。我们发现这一不一致源于极小扩散时间步下得分模型的误差,此时模型需准确捕捉数据分布的演化。在此区间,扩散模型无法满足控制得分演化的福克-普朗克方程。我们将此偏差视为不一致的主要来源,并提出一种基于福克-普朗克方程的正则化能量扩散模型以强制一致性。我们在多种生物分子系统(包括快速折叠蛋白)上验证了该方法的有效性,引入了一种先进的可迁移玻尔兹曼模拟器,支持高效模拟与采样,显著提升一致性。代码、模型权重及自包含的JAX和PyTorch笔记本已开源。
原文摘要 · Abstract (English)
In recent years, diffusion models trained on equilibrium molecular distributions have proven effective for sampling biomolecules. Beyond direct sampling, the score of such a model can also be used to derive the forces that act on molecular systems. However, while classical diffusion sampling usually recovers the training distribution, the corresponding energy-based interpretation of the learned score is often inconsistent with this distribution, even for low-dimensional toy systems. We trace this inconsistency to inaccuracies of the learned score at very small diffusion timesteps, where the model must capture the correct evolution of the data distribution. In this regime, diffusion models fail to satisfy the Fokker-Planck equation, which governs the evolution of the score. We interpret this deviation as one source of the observed inconsistencies and propose an energy-based diffusion model with a Fokker-Planck-derived regularization term to enforce consistency. We demonstrate our approach by sampling and simulating multiple biomolecular systems, including fast-folding proteins, and by introducing a state-of-the-art transferable Boltzmann emulator for dipeptides that supports simulation and achieves improved consistency and efficient sampling. Our code, model weights, and self-contained JAX and PyTorch notebooks are available at https://github.com/noegroup/ScoreMD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。