用结构反馈优化蛋白序列设计,让生成的氨基酸序列更易折叠成目标结构。
Protein Inverse Folding From Structure Feedback
- 通过结构预测模型生成偏好标签,用偏好优化方法微调序列设计模型。
- 在CATH数据集上平均TM-Score从0.77提升至0.81,结构相似性显著增强。
- 对难设计结构迭代优化,平均性能提升79.5%,适合蛋白质工程研究者。
逆折叠问题旨在设计能折叠成目标三维结构的氨基酸序列,对生物技术应用至关重要。本文提出一种新方法,利用直接偏好优化(DPO)微调逆折叠模型,反馈来自蛋白质折叠模型。给定目标结构后,先从逆折叠模型采样候选序列,再用折叠模型预测其三维结构,生成成对结构偏好标签,用于DPO目标下的微调。在CATH 4.2测试集上的结果表明,该方法不仅提升了基线模型的序列恢复能力,还将平均TM-Score从0.77提高到0.81,表明结构相似性显著增强。此外,在挑战性蛋白结构上迭代应用该方法,平均TM-Score相较基线提升79.5%。本工作为利用结构反馈有效提升蛋白序列设计能力提供了新方向。
原文摘要 · Abstract (English)
The inverse folding problem, aiming to design amino acid sequences that fold into desired three-dimensional structures, is pivotal for various biotechnological applications. Here, we introduce a novel approach leveraging Direct Preference Optimization (DPO) to fine-tune an inverse folding model using feedback from a protein folding model. Given a target protein structure, we begin by sampling candidate sequences from the inverse-folding model, then predict the three-dimensional structure of each sequence with the folding model to generate pairwise structural-preference labels. These labels are used to fine-tune the inverse-folding model under the DPO objective. Our results on the CATH 4.2 test set demonstrate that DPO fine-tuning not only improves sequence recovery of baseline models but also leads to a significant improvement in average TM-Score from 0.77 to 0.81, indicating enhanced structure similarity. Furthermore, iterative application of our DPO-based method on challenging protein structures yields substantial gains, with an average TM-Score increase of 79.5\% with regard to the baseline model. This work establishes a promising direction for enhancing protein sequence design ability from structure feedback by effectively utilizing preference optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。