arXiv:2410.19222cs.LGq-bio.QM2024-10被引 14

用预训练语言模型生成具有特定功能的肽,提升合成生物学设计效率。

Peptide-GPT: Generative Design of Peptides using Generative Pre-trained Transformers and Bio-informatic Supervision

  • 基于NLP的肽序列生成模型,结合生物信息学筛选确保结构合理性。
  • 在溶血、非溶血、抗污和溶解性任务上准确率分别达76.26%~78.84%。
  • 适合从事蛋白质设计、药物开发与合成生物学的研究者使用。

近年来,自然语言处理模型在文本生成以外的多个领域展现出强大能力。本文提出PeptideGPT,一种专用于生成具有特定性质(溶血性、溶解性、抗污性)的蛋白序列的蛋白质语言模型。为严格评估生成序列,我们构建了包含生物信息学方法的综合评估流程:首先根据困惑度得分对生成序列排序,再通过蛋白质凸包过滤排除异常序列,最后使用ESMFold预测结构,并筛选pLDDT值大于70的序列以保证有序结构。生成序列的性质由任务专用分类器PeptideBERT和HAPPENN评估,结果显示在溶血性、非溶血性、抗污性和溶解性生成任务上的准确率分别为76.26%、72.46%、78.84%和68.06%。实验表明PeptideGPT在从头蛋白质设计中有效,凸显了基于NLP方法在合成生物学与生物信息学中的应用潜力。代码、模型与数据已开源:https://github.com/aayush-shah14/PeptideGPT。

原文摘要 · Abstract (English)

In recent years, natural language processing (NLP) models have demonstrated remarkable capabilities in various domains beyond traditional text generation. In this work, we introduce PeptideGPT, a protein language model tailored to generate protein sequences with distinct properties: hemolytic activity, solubility, and non-fouling characteristics. To facilitate a rigorous evaluation of these generated sequences, we established a comprehensive evaluation pipeline consisting of ideas from bioinformatics to retain valid proteins with ordered structures. First, we rank the generated sequences based on their perplexity scores, then we filter out those lying outside the permissible convex hull of proteins. Finally, we predict the structure using ESMFold and select the proteins with pLDDT values greater than 70 to ensure ordered structure. The properties of generated sequences are evaluated using task-specific classifiers - PeptideBERT and HAPPENN. We achieved an accuracy of 76.26% in hemolytic, 72.46% in non-hemolytic, 78.84% in non-fouling, and 68.06% in solubility protein generation. Our experimental results demonstrate the effectiveness of PeptideGPT in de novo protein design and underscore the potential of leveraging NLP-based approaches for paving the way for future innovations and breakthroughs in synthetic biology and bioinformatics. Codes, models, and data used in this study are freely available at: https://github.com/aayush-shah14/PeptideGPT.

蛋白质设计生成模型生物信息学肽序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。