用强化学习提升肽设计多样性,让序列更丰富且结构更稳定。
Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization
- 通过直接偏好优化+在线多样性正则化改进蛋白质折叠模型
- 肽序列多样性提升20%,结构相似度仍优于基线模型8%以上
- 适合需要多样化肽序列的药物与材料设计研究者
逆折叠模型在基于结构的设计中至关重要,可预测折叠为指定参考结构的氨基酸序列。如ProteinMPNN这类消息传递编码器-解码器模型,能可靠地从参考结构生成新序列。然而,在肽类设计中,这些模型易生成重复且无法正确折叠的序列。为此,我们采用直接偏好优化(DPO)微调ProteinMPNN,以生成多样且结构一致的肽序列。提出两项改进:在线多样性正则化与领域特定先验。此外,我们对提升解码器模型多样性有了新认识。在OpenFold生成的结构条件下,微调后模型达到当前最优结构相似度,较基线ProteinMPNN提升至少8%。相较于标准DPO,我们的方法在不损失结构相似度的前提下,序列多样性最高提升20%。
原文摘要 · Abstract (English)
Inverse folding models play an important role in structure-based design by predicting amino acid sequences that fold into desired reference structures. Models like ProteinMPNN, a message-passing encoder-decoder model, are trained to reliably produce new sequences from a reference structure. However, when applied to peptides, these models are prone to generating repetitive sequences that do not fold into the reference structure. To address this, we fine-tune ProteinMPNN to produce diverse and structurally consistent peptide sequences via Direct Preference Optimization (DPO). We derive two enhancements to DPO: online diversity regularization and domain-specific priors. Additionally, we develop a new understanding on improving diversity in decoder models. When conditioned on OpenFold generated structures, our fine-tuned models achieve state-of-the-art structural similarity scores, improving base ProteinMPNN by at least 8%. Compared to standard DPO, our regularized method achieves up to 20% higher sequence diversity with no loss in structural similarity score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。