用图神经网络与扩散模型生成完整蛋白结构,精准还原动态构象。
Generative Modeling of Full-Atom Protein Conformations using Latent Diffusion on Graph Embeddings
- 用图嵌入+扩散模型在隐空间生成蛋白构象,保留所有原子细节。
- 在12000帧的多尺度分子动力学数据上,全原子结构精度达0.7(lDDT)。
- 适合药物设计领域研究复杂动态蛋白靶点,开源数据与代码可用。
生成如G蛋白偶联受体(GPCRs)等动态蛋白的多样化、全原子构象集合对理解其功能至关重要,但现有生成模型常简化原子细节或忽略构象多样性。本文提出潜变量扩散蛋白生成框架(LD-FPG),直接从分子动力学(MD)轨迹生成完整全原子蛋白结构,包含每个侧链重原子。该方法采用切比雪夫图神经网络(ChebNet)获取蛋白构象的低维隐空间表示,并结合三种池化策略:盲池化、顺序池化和残基池化。在这些隐表示上训练扩散模型,生成新样本后通过解码器映射回笛卡尔坐标,可选地引入二面角损失进行正则化。基于人类多巴胺D2受体在膜环境中的2微秒MD轨迹(12,000帧,D2R-MD),顺序与残基池化策略以高结构保真度重现参考构象集(全原子lDDT约0.7;α碳lDDT约0.8),并使主链与侧链二面角分布的杰恩-申诺尔散度小于0.03,与原始MD数据高度一致。该方法为大蛋白提供系统特异性全原子构象集生成路径,是复杂动态靶点结构药物设计的有力工具。相关数据集与实现已公开,便于后续研究。
原文摘要 · Abstract (English)
Generating diverse, all-atom conformational ensembles of dynamic proteins such as G-protein-coupled receptors (GPCRs) is critical for understanding their function, yet most generative models simplify atomic detail or ignore conformational diversity altogether. We present latent diffusion for full protein generation (LD-FPG), a framework that constructs complete all-atom protein structures, including every side-chain heavy atom, directly from molecular dynamics (MD) trajectories. LD-FPG employs a Chebyshev graph neural network (ChebNet) to obtain low-dimensional latent embeddings of protein conformations, which are processed using three pooling strategies: blind, sequential and residue-based. A diffusion model trained on these latent representations generates new samples that a decoder, optionally regularized by dihedral-angle losses, maps back to Cartesian coordinates. Using D2R-MD, a 2-microsecond MD trajectory (12 000 frames) of the human dopamine D2 receptor in a membrane environment, the sequential and residue-based pooling strategy reproduces the reference ensemble with high structural fidelity (all-atom lDDT of approximately 0.7; C-alpha-lDDT of approximately 0.8) and recovers backbone and side-chain dihedral-angle distributions with a Jensen-Shannon divergence of less than 0.03 compared to the MD data. LD-FPG thereby offers a practical route to system-specific, all-atom ensemble generation for large proteins, providing a promising tool for structure-based therapeutic design on complex, dynamic targets. The D2R-MD dataset and our implementation are freely available to facilitate further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。