arXiv:2502.15610cs.LGcs.AI2025-02被引 2

PDeepPP用统一模型精准识别多种肽功能,助力新药研发。

A general language model for peptide function identification

  • 融合预训练蛋白语言模型与混合架构,捕捉肽序列全局与局部特征。
  • 在33项任务中25项达顶尖水平,抗菌与磷酸化位点识别准确率超0.97。
  • 开源模型代码数据,适合药物发现与功能注释研究者使用。

准确识别生物活性肽(BPs)和蛋白质翻译后修饰(PTMs)对理解蛋白功能和推动治疗发现至关重要。然而,现有计算方法在跨不同肽功能时泛化能力有限。本文提出PDeepPP,一种整合预训练蛋白语言模型与混合Transformer-CNN架构的统一深度学习框架,可鲁棒识别多种肽类及PTM位点。我们构建了全面基准数据集,并采用策略缓解数据不平衡问题,使PDeepPP能系统提取全局与局部序列特征。通过降维分析与对比实验,PDeepPP展现出强解释性表征,在33项生物识别任务中25项达到领先性能。显著成果包括:抗菌肽识别准确率0.9726,磷酸化位点识别准确率0.9984,糖基化位点预测特异性达99.5%,抗疟疾任务假阴性显著降低。该模型支持大规模、高精度肽分析,助力生物医学研究与疾病治疗新靶点发现。所有代码、数据集及预训练模型已公开于GitHub(https://github.com/fondress/PDeepPP)与Hugging Face(https://huggingface.co/fondress/PDeppPP)。

原文摘要 · Abstract (English)

Accurate identification of bioactive peptides (BPs) and protein post-translational modifications (PTMs) is essential for understanding protein function and advancing therapeutic discovery. However, most computational methods remain limited in their generalizability across diverse peptide functions. Here, we present PDeepPP, a unified deep learning framework that integrates pretrained protein language models with a hybrid transformer-CNN architecture, enabling robust identification across diverse peptide classes and PTM sites. We curated comprehensive benchmark datasets and implemented strategies to address data imbalance, allowing PDeepPP to systematically extract both global and local sequence features. Through extensive analyses including dimensionality reduction and comparison studies, PDeepPP demonstrates strong, interpretable peptide representations and achieves state-of-the-art performance in 25 of the 33 biological identification tasks. Notably, PDeepPP attains high accuracy in antimicrobial (0.9726) and phosphorylation site (0.9984) identification, with 99.5% specificity in glycosylation site prediction and substantial reduction in false negatives in antimalarial tasks. By enabling large-scale, accurate peptide analysis, PDeepPP supports biomedical research and the discovery of novel therapeutic targets for disease treatment. All code, datasets, and pretrained models are publicly available via GitHub (https://github.com/fondress/PDeepPP) and Hugging Face (https://huggingface.co/fondress/PDeppPP)

肽功能识别深度学习药物发现预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。