用扩散模型+特征对齐,提升蛋白质逆折叠序列预测精度
Diffusion Model with Representation Alignment for Protein Inverse Folding
- 引入共享中心聚合全局结构信息,按需分配给残基
- 在去噪过程中对齐噪声表示与语义表示,提升预测准确性
- 在多个数据集上表现领先,适合蛋白质设计研究者
蛋白质逆折叠是生物信息学中的基础问题,旨在从给定的蛋白质骨架结构恢复氨基酸序列。尽管现有方法取得一定进展,但仍难以充分捕捉残基间的复杂关系,影响序列预测精度。本文提出一种基于扩散模型与表征对齐(DMRA)的新方法:(1) 设计一个共享中心,聚合整个蛋白结构的上下文信息,并选择性地分发给每个残基;(2) 在去噪过程中,将带噪隐藏表示与干净语义表示对齐。通过预定义的氨基酸类型语义表示,利用类型嵌入作为语义反馈,对每个残基进行归一化。在CATH4.2数据集上进行大量实验表明,DMRA优于现有先进方法,达到当前最优性能,并在TS50和TS500数据集上展现出强大的泛化能力。
原文摘要 · Abstract (English)
Protein inverse folding is a fundamental problem in bioinformatics, aiming to recover the amino acid sequences from a given protein backbone structure. Despite the success of existing methods, they struggle to fully capture the intricate inter-residue relationships critical for accurate sequence prediction. We propose a novel method that leverages diffusion models with representation alignment (DMRA), which enhances diffusion-based inverse folding by (1) proposing a shared center that aggregates contextual information from the entire protein structure and selectively distributes it to each residue; and (2) aligning noisy hidden representations with clean semantic representations during the denoising process. This is achieved by predefined semantic representations for amino acid types and a representation alignment method that utilizes type embeddings as semantic feedback to normalize each residue. In experiments, we conduct extensive evaluations on the CATH4.2 dataset to demonstrate that DMRA outperforms leading methods, achieving state-of-the-art performance and exhibiting strong generalization capabilities on the TS50 and TS500 datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。