用扩散模型同时设计蛋白结合剂的序列与结构,无需反复实验验证。
PPDiff: Diffusing in Hybrid Sequence-Structure Space for Protein-Protein Complex Design

- 采用序列-结构交错网络,融合注意力与图神经网络捕捉远近关系。
- 在70多万个复合物数据上预训练,下游任务成功率最高达50%。
- 适合需要快速设计高亲和力蛋白结合剂的研究者使用。
设计高亲和力的蛋白结合蛋白对生物医学研究与生物技术至关重要。尽管已有针对特定蛋白的进展,但无需大量湿实验即可按需设计任意靶点的高亲和力结合剂,仍是重大挑战。本文提出PPDiff,一种非自回归扩散模型,可联合设计任意蛋白靶点的结合剂序列与三维结构。该模型基于我们提出的序列-结构交错网络(SSINC),集成交错自注意力层以捕获全局氨基酸关联,k最近邻(kNN)等变图层以建模3D空间中的局部相互作用,以及因果注意力层以简化序列内部复杂依赖。为评估模型,我们构建了包含706,360个复合物的通用蛋白-蛋白复合物数据集PPBench。PPDiff在该数据集上预训练,并在两个真实应用场景中微调:靶向蛋白小结合剂设计与抗原-抗体复合物设计。模型表现持续优于基线方法,在预训练任务与两项下游任务中成功率分别达到50.00%、23.16%和16.89%。代码、数据与模型已开源于https://github.com/JocelynSong/PPDiff。
原文摘要 · Abstract (English)
Designing protein-binding proteins with high affinity is critical in biomedical research and biotechnology. Despite recent advancements targeting specific proteins, the ability to create high-affinity binders for arbitrary protein targets on demand, without extensive rounds of wet-lab testing, remains a significant challenge. Here, we introduce PPDiff, a diffusion model to jointly design the sequence and structure of binders for arbitrary protein targets in a non-autoregressive manner. PPDiffbuilds upon our developed Sequence Structure Interleaving Network with Causal attention layers (SSINC), which integrates interleaved self-attention layers to capture global amino acid correlations, k-nearest neighbor (kNN) equivariant graph layers to model local interactions in three-dimensional (3D) space, and causal attention layers to simplify the intricate interdependencies within the protein sequence. To assess PPDiff, we curate PPBench, a general protein-protein complex dataset comprising 706,360 complexes from the Protein Data Bank (PDB). The model is pretrained on PPBenchand finetuned on two real-world applications: target-protein mini-binder complex design and antigen-antibody complex design. PPDiffconsistently surpasses baseline methods, achieving success rates of 50.00%, 23.16%, and 16.89% for the pretraining task and the two downstream applications, respectively. The code, data and models are available at https://github.com/JocelynSong/PPDiff.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。