用结构偏好优化提升蛋白质序列设计成功率。
Improving Protein Sequence Design through Designability Preference Optimization
- 以AlphaFold pLDDT为信号,通过偏好优化改进生成目标。
- 残基级优化使设计成功率从6.56%升至17.57%。
- 适合需要高折叠成功率的蛋白质设计研究者。
蛋白质序列设计方法在从头设计中表现优异,但其训练目标为序列恢复,无法保证设计可折叠性——即设计序列正确折叠成目标结构的概率。为此,本文重新定义训练目标,引导序列生成向高可设计性方向优化。我们引入直接偏好优化(DPO),以AlphaFold pLDDT分数作为偏好信号,显著提升了体外设计成功率。为进一步实现残基级别的精细化优化,提出残基级设计可折叠性偏好优化(ResiDPO),采用残基级结构奖励并解耦各残基优化过程,可在不破坏已有良好区域的前提下直接提升设计可折叠性。基于带有残基级标注的精选数据集,我们对LigandMPNN进行ResiDPO微调,得到EnhancedMPNN,在具有挑战性的酶设计基准上,体外设计成功率接近提升3倍(从6.56%增至17.57%)。
原文摘要 · Abstract (English)
Protein sequence design methods have demonstrated strong performance in sequence generation for de novo protein design. However, as the training objective was sequence recovery, it does not guarantee designability--the likelihood that a designed sequence folds into the desired structure. To bridge this gap, we redefine the training objective by steering sequence generation toward high designability. To do this, we integrate Direct Preference Optimization (DPO), using AlphaFold pLDDT scores as the preference signal, which significantly improves the in silico design success rate. To further refine sequence generation at a finer, residue-level granularity, we introduce Residue-level Designability Preference Optimization (ResiDPO), which applies residue-level structural rewards and decouples optimization across residues. This enables direct improvement in designability while preserving regions that already perform well. Using a curated dataset with residue-level annotations, we fine-tune LigandMPNN with ResiDPO to obtain EnhancedMPNN, which achieves a nearly 3-fold increase in in silico design success rate (from 6.56% to 17.57%) on a challenging enzyme design benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。