用检索增强扩散模型设计蛋白质序列,准确率提升19%
RadDiff: Retrieval-Augmented Denoising Diffusion for Protein Inverse Folding
- 引入检索增强机制获取最新蛋白知识
- 扩散过程融合外部知识,序列恢复率最高提升19%
- 适合需要高可折叠性序列的设计场景
蛋白质逆折叠——根据目标结构设计氨基酸序列——是计算蛋白质工程的核心问题。现有方法或不利用外部知识,或依赖蛋白语言模型(PLMs),前者忽略自然蛋白数据中的知识,后者参数效率低且难以适应不断增长的蛋白数据。为此,本文提出一种新方法RadDiff(检索增强去噪扩散),设计了新颖的检索增强机制以捕获最新蛋白知识,并构建轻量级知识感知扩散模型,将知识融入扩散过程。在CATH、TS50和PDB2022数据集上的实验表明,RadDiff持续优于现有方法,序列恢复率最高提升19%。结果还显示,RadDiff生成的序列具有高可折叠性,且能随数据库规模有效扩展。
原文摘要 · Abstract (English)
Protein inverse folding, the design of an amino acid sequence based on a target protein structure, is a fundamental problem of computational protein engineering. Existing methods either generate sequences without leveraging external knowledge or relying on protein language models~(PLMs). The former omits the knowledge stored in natural protein data, while the latter is parameter-inefficient and inflexible to adapt to ever-growing protein data. To overcome the above drawbacks, in this paper we propose a novel method, called $\underline{\text{r}}$etrieval-$\underline{\text{a}}$ugmented $\underline{\text{d}}$enoising $\underline{\text{diff}}$usion~($\mbox{RadDiff}$), for protein inverse folding. In RadDiff, a novel retrieval-augmentation mechanism is designed to capture the up-to-date protein knowledge. We further design a knowledge-aware diffusion model that integrates this protein knowledge into the diffusion process via a lightweight module. Experimental results on the CATH, TS50, and PDB2022 datasets show that $\mbox{RadDiff}$ consistently outperforms existing methods, improving sequence recovery rate by up to 19\%. Experimental results also demonstrate that RadDiff generates highly foldable sequences and scales effectively with database size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。