用机器学习力场加速分子构象生成,训练一次即可高效生成稳定结构。
Elign: Equivariant Diffusion Model Alignment from Foundational Machine Learning Force Fields
- 用预训练的机器学习力场替代耗时的量子计算,提供物理引导信号
- 通过强化学习在训练阶段优化生成策略,使采样时无需额外计算能量
- 生成的分子构象更接近真实能量最低态,适合需要高质量构象的药物设计
生成三维分子构象的生成模型必须遵守欧几里得对称性,并将概率质量集中在热力学有利、机械稳定的结构上。然而,现有的E(3)等变扩散模型往往继承半经验训练数据的偏差,而非反映高保真哈密顿量的平衡分布。尽管基于物理的引导可纠正此问题,但面临两个计算瓶颈:昂贵的量子化学评估(如DFT)以及每步采样均需重复查询。我们提出Elign,一种后训练框架,以摊销这两项成本。首先,用更快的预训练基础机器学习力场(MLFF)替代昂贵的DFT评估,提供物理信号;其次,通过将物理引导转移到训练阶段,消除运行时重复查询。为实现第二项摊销,我们将反向扩散建模为强化学习问题,引入力-能解耦组相对策略优化(FED-GRPO),独立优化基于势能的能量奖励与基于力的稳定性奖励。实验表明,Elign生成的构象具有更低的金标准DFT能量和力,同时提升稳定性。关键在于推理速度与无引导采样相同,因生成过程无需能量评估。
原文摘要 · Abstract (English)
Generative models for 3D molecular conformations must respect Euclidean symmetries and concentrate probability mass on thermodynamically favorable, mechanically stable structures. However, E(3)-equivariant diffusion models often reproduce biases from semi-empirical training data rather than capturing the equilibrium distribution of a high-fidelity Hamiltonian. While physics-based guidance can correct this, it faces two computational bottlenecks: expensive quantum-chemical evaluations (e.g., DFT) and the need to repeat such queries at every sampling step. We present Elign, a post-training framework that amortizes both costs. First, we replace expensive DFT evaluations with a faster, pretrained foundational machine-learning force field (MLFF) to provide physical signals. Second, we eliminate repeated run-time queries by shifting physical steering to the training phase. To achieve the second amortization, we formulate reverse diffusion as a reinforcement learning problem and introduce Force--Energy Disentangled Group Relative Policy Optimization (FED-GRPO) to fine-tune the denoising policy. FED-GRPO includes a potential-based energy reward and a force-based stability reward, which are optimized and group-normalized independently. Experiments show that Elign generates conformations with lower gold-standard DFT energies and forces, while improving stability. Crucially, inference remains as fast as unguided sampling, since no energy evaluations are required during generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。