用掩码先验引导扩散模型,提升无序区等难预测残基的蛋白质逆折叠生成精度。
Mask prior-guided denoising diffusion improves inverse protein folding
- 基于图结构网络和掩码先验预训练,融合结构信息与残基相互作用。
- 在4个基准测试中显著超越当前最优方法,生成序列接近天然蛋白特性。
- 适合需要高保真度蛋白质序列设计的研究者,尤其关注结构不确定性区域。
逆蛋白质折叠旨在生成能折叠成目标结构的有效氨基酸序列,近年深度学习取得显著进展并展现出竞争力。然而,仍面临高结构不确定区域(如无序区)预测困难等问题。为此,我们提出一种掩码先验引导的去噪扩散框架MapDiff,可精准捕捉结构信息与残基相互作用。MapDiff是一种离散扩散概率模型,通过迭代去噪生成条件于给定蛋白骨架的氨基酸序列。为融入结构信息与残基交互,我们设计了基于图的去噪网络,并采用掩码先验预训练策略。生成过程中,结合去噪扩散隐式模型与蒙特卡洛丢弃以降低不确定性。在四个具有挑战性的序列设计基准上评估显示,MapDiff显著优于现有最先进方法。此外,MapDiff生成的体外序列在不同蛋白质家族与架构中,其物理化学与结构特征均高度接近天然蛋白。
原文摘要 · Abstract (English)
Inverse protein folding generates valid amino acid sequences that can fold into a desired protein structure, with recent deep-learning advances showing strong potential and competitive performance. However, challenges remain, such as predicting elements with high structural uncertainty, including disordered regions. To tackle such low-confidence residue prediction, we propose a Mask-prior-guided denoising Diffusion (MapDiff) framework that accurately captures both structural information and residue interactions for inverse protein folding. MapDiff is a discrete diffusion probabilistic model that iteratively generates amino acid sequences with reduced noise, conditioned on a given protein backbone. To incorporate structural information and residue interactions, we develop a graph-based denoising network with a mask-prior pre-training strategy. Moreover, in the generative process, we combine the denoising diffusion implicit model with Monte-Carlo dropout to reduce uncertainty. Evaluation on four challenging sequence design benchmarks shows that MapDiff substantially outperforms state-of-the-art methods. Furthermore, the in silico sequences generated by MapDiff closely resemble the physico-chemical and structural characteristics of native proteins across different protein families and architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。