用结构检索提升抗体序列设计准确率,避免生成无效序列。
Fast and Accurate Antibody Sequence Design via Structure Retrieval
- 通过检索天然抗体数据库中的相似结构来推断抗体序列。
- 在抗体和T细胞受体序列恢复上优于现有方法。
- 适合需要高精度抗体设计的药物研发人员。
近年来蛋白质设计利用扩散模型生成结构骨架,再通过蛋白逆折叠推断序列。但针对抗体互补决定区(CDRs)等高度可变结构,传统方法常因幻觉导致非功能序列。本文提出Igseek,一种基于结构检索的新框架:先用多通道等变图神经网络生成高质的CDR骨架几何表示,再从天然抗体数据库中检索结构相似片段,对齐序列并利用保守序列基元提升推断准确性。实验表明,Igseek在结构检索效率和抗体、T细胞受体序列恢复上均优于当前最优方法,为治疗性蛋白质设计提供了新的检索驱动视角。
原文摘要 · Abstract (English)
Recent advancements in protein design have leveraged diffusion models to generate structural scaffolds, followed by a process known as protein inverse folding, which involves sequence inference on these scaffolds. However, these methodologies face significant challenges when applied to hyper-variable structures such as antibody Complementarity-Determining Regions (CDRs), where sequence inference frequently results in non-functional sequences due to hallucinations. Distinguished from prevailing protein inverse folding approaches, this paper introduces Igseek, a novel structure-retrieval framework that infers CDR sequences by retrieving similar structures from a natural antibody database. Specifically, Igseek employs a simple yet effective multi-channel equivariant graph neural network to generate high-quality geometric representations of CDR backbone structures. Subsequently, it aligns sequences of structurally similar CDRs and utilizes structurally conserved sequence motifs to enhance inference accuracy. Our experiments demonstrate that Igseek not only proves to be highly efficient in structural retrieval but also outperforms state-of-the-art approaches in sequence recovery for both antibodies and T-Cell Receptors, offering a new retrieval-based perspective for therapeutic protein design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。