arXiv:2602.20198q-bio.QMcs.LG2026-02被引 1

融合深度蛋白嵌入与手工特征,提升促炎肽预测准确率

KEMP-PIP: A Feature-Fusion Based Approach for Pro-inflammatory Peptide Prediction

  • 结合ESM语言模型嵌入与多尺度氨基酸组合特征
  • 在标准数据集上达MCC 0.505,优于现有方法9.5%以上
  • 适合免疫学研究者快速筛选潜在促炎肽

促炎肽(PIPs)在免疫信号传导和炎症中起关键作用,但因实验成本高、耗时长,难以通过实验识别。为此,我们提出KEMP-PIP,一种融合深度蛋白嵌入与手工描述符的混合机器学习框架。该方法结合预训练ESM蛋白语言模型的上下文嵌入、多尺度k-mer频率、理化性质描述符及modlAMP序列特征。通过特征剪枝与类别加权逻辑回归处理高维与类别不平衡问题,采用集成平均与优化决策阈值提升敏感性-特异性平衡。系统消融实验证明,互补特征集的整合持续提升预测性能。在标准基准数据集上,KEMP-PIP取得MCC 0.505、准确率0.752、AUC 0.762,优于ProIn-fuse、MultiFeatVotPIP和StackPIP。相较StackPIP,MCC提升9.5%,准确率与AUC各提升4.8%。KEMP-PIP在线服务器免费开放(https://nilsparrow1920-kemp-pip.hf.space/),完整代码见GitHub(https://github.com/S18-Niloy/KEMP-PIP)。

原文摘要 · Abstract (English)

Pro-inflammatory peptides (PIPs) play critical roles in immune signaling and inflammation but are difficult to identify experimentally due to costly and time-consuming assays. To address this challenge, we present KEMP-PIP, a hybrid machine learning framework that integrates deep protein embeddings with handcrafted descriptors for robust PIP prediction. Our approach combines contextual embeddings from pretrained ESM protein language models with multi-scale k-mer frequencies, physicochemical descriptors, and modlAMP sequence features. Feature pruning and class-weighted logistic regression manage high dimensionality and class imbalance, while ensemble averaging with an optimized decision threshold enhances the sensitivity--specificity balance. Through systematic ablation studies, we demonstrate that integrating complementary feature sets consistently improves predictive performance. On the standard benchmark dataset, KEMP-PIP achieves an MCC of 0.505, accuracy of 0.752, and AUC of 0.762, outperforming ProIn-fuse, MultiFeatVotPIP, and StackPIP. Relative to StackPIP, these results represent improvements of 9.5% in MCC and 4.8% in both accuracy and AUC. The KEMP-PIP web server is freely available at https://nilsparrow1920-kemp-pip.hf.space/ and the full implementation at https://github.com/S18-Niloy/KEMP-PIP.

肽预测机器学习免疫学蛋白质语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。