用势能梯度引导生成模型,让分子构象采样更准更快。
Potential Score Matching: Debiasing Molecular Structure Sampling with Potential Energy Guidance
- 利用势能梯度指导生成模型,无需精确能量函数。
- 在LJ势能模型上优于现有顶尖方法,采样分布更接近玻尔兹曼分布。
- 适合需要高效精准分子构象采样的化学与材料研究者。
分子物理性质的系综平均与其构象分布密切相关,而准确采样该分布是物理与化学中的基础挑战。传统方法如分子动力学(MD)和马尔可夫链蒙特卡洛(MCMC)耗时且成本高。近期扩散模型因其高效性成为替代方案,但获取无偏目标分布仍昂贵,主要因需满足遍历性。为此,本文提出势能得分匹配(PSM),利用势能梯度引导生成模型。该方法无需精确能量函数,即使在有限且有偏差的数据上训练,也能有效消除采样偏差。在常用简化模型Lennard-Jones(LJ)势能上,PSM性能超越现有最先进(SOTA)模型。进一步在高维问题上使用MD17和MD22数据集评估,结果表明PSM生成的分子分布更接近玻尔兹曼分布,优于传统扩散模型。
原文摘要 · Abstract (English)
The ensemble average of physical properties of molecules is closely related to the distribution of molecular conformations, and sampling such distributions is a fundamental challenge in physics and chemistry. Traditional methods like molecular dynamics (MD) simulations and Markov chain Monte Carlo (MCMC) sampling are commonly used but can be time-consuming and costly. Recently, diffusion models have emerged as efficient alternatives by learning the distribution of training data. Obtaining an unbiased target distribution is still an expensive task, primarily because it requires satisfying ergodicity. To tackle these challenges, we propose Potential Score Matching (PSM), an approach that utilizes the potential energy gradient to guide generative models. PSM does not require exact energy functions and can debias sample distributions even when trained on limited and biased data. Our method outperforms existing state-of-the-art (SOTA) models on the Lennard-Jones (LJ) potential, a commonly used toy model. Furthermore, we extend the evaluation of PSM to high-dimensional problems using the MD17 and MD22 datasets. The results demonstrate that molecular distributions generated by PSM more closely approximate the Boltzmann distribution compared to traditional diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。