用学习的隐空间模拟蛋白质原子级动态,突破传统方法局限。
Beyond Ensembles: Simulating All-Atom Protein Dynamics in a Learned Latent Space
- 在学习的隐空间中构建动态传播器,实现原子级构象演化模拟。
- 自回归神经网络表现最优,能保持长期稳定与物理时间尺度。
- 适用于大分子、复杂构象转换的精准动力学建模,适合生物学家与结构科学家。
模拟生物分子的长时序动态是计算科学的核心挑战。尽管增强采样方法可加速模拟,但依赖预定义集体变量,难以捕捉复杂稳态间的切换机制。近期生成模型LD-FPG通过从参考结构学习所有原子形变来生成静态平衡系综,为全原子构象生成提供强大工具。然而该方法未建模构象间的演化过程。本文引入图隐空间动力学传播器(GLDP),并比较三类传播器:(i)基于得分的Langevin动力学,(ii)Koopman线性算子,(iii)自回归神经网络。在统一的编码-传播-解码框架下,评估了长时间稳定性、主链与侧链系综保真度及时间动力学(通过TICA)。对小肽、混合拓扑蛋白及大型G蛋白偶联受体的基准测试表明,自回归神经网络在长期滚动中表现最稳健,且具有连贯的物理时间尺度;得分引导的Langevin在得分学习良好时最佳恢复侧链热力学;Koopman提供轻量可解释基线,但倾向于抑制波动。结果厘清了传播器间权衡,为全原子蛋白动态隐空间模拟提供了实用指导。
原文摘要 · Abstract (English)
Simulating the long-timescale dynamics of biomolecules is a central challenge in computational science. While enhanced sampling methods can accelerate these simulations, they rely on pre-defined collective variables that are often difficult to identify, restricting their ability to model complex switching mechanisms between metastable states. A recent generative model, LD-FPG, demonstrated that this problem could be bypassed by learning to sample the static equilibrium ensemble as all-atom deformations from a reference structure, establishing a powerful method for all-atom ensemble generation. However, while this approach successfully captures a system's probable conformations, it does not model the temporal evolution between them. We introduce the Graph Latent Dynamics Propagator (GLDP), a modular component for simulating dynamics within the learned latent space of LD-FPG. We then compare three classes of propagators: (i) score-guided Langevin dynamics, (ii) Koopman-based linear operators, and (iii) autoregressive neural networks. Within a unified encoder-propagator-decoder framework, we evaluate long-horizon stability, backbone and side-chain ensemble fidelity, and temporal kinetics via TICA. Benchmarks on systems ranging from small peptides to mixed-topology proteins and large GPCRs reveal that autoregressive neural networks deliver the most robust long rollouts and coherent physical timescales; score-guided Langevin best recovers side-chain thermodynamics when the score is well learned; and Koopman provides an interpretable, lightweight baseline that tends to damp fluctuations. These results clarify the trade-offs among propagators and offer practical guidance for latent-space simulators of all-atom protein dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。