用预训练蛋白大模型预测淀粉样蛋白区域,准确率达84.5%。
Deep Learning Model for Amyloidogenicity Prediction using a Pre-trained Protein LLM
- 基于预训练蛋白LLM提取序列上下文特征,结合双向LSTM与GRU建模。
- 10折交叉验证准确率84.5%,测试集准确率达83%。
- 适合蛋白质功能预测与阿尔茨海默病研究者参考。
肽和蛋白质的淀粉样蛋白形成能力预测仍是生物信息学的重要课题。当前方法多依赖进化模式和氨基酸个体属性,而基于序列信息的特征展现出了优异的预测性能。本研究利用预训练蛋白大语言模型提取蛋白质序列的上下文特征,结合双向LSTM与GRU网络,预测肽和蛋白质序列中的淀粉样蛋白区域。在10折交叉验证中达到84.5%的准确率,在独立测试集上为83%。结果表明该方法具有竞争力,凸显了大语言模型在提升淀粉样蛋白预测精度方面的潜力。
原文摘要 · Abstract (English)
The prediction of amyloidogenicity in peptides and proteins remains a focal point of ongoing bioinformatics. The crucial step in this field is to apply advanced computational methodologies. Many recent approaches to predicting amyloidogenicity within proteins are highly based on evolutionary motifs and the individual properties of amino acids. It is becoming increasingly evident that the sequence information-based features show high predictive performance. Consequently, our study evaluated the contextual features of protein sequences obtained from a pretrained protein large language model leveraging bidirectional LSTM and GRU to predict amyloidogenic regions in peptide and protein sequences. Our method achieved an accuracy of 84.5% on 10-fold cross-validation and an accuracy of 83% in the test dataset. Our results demonstrate competitive performance, highlighting the potential of LLMs in enhancing the accuracy of amyloid prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。