用自监督+微调两阶段模型提升蛋白质功能预测精度
Protein-Mamba: Biological Mamba Models for Protein Function Prediction
- 分两阶段:先自监督学通用结构,再用标注数据微调
- 在多个蛋白功能数据集上表现优于主流方法
- 适合药物发现领域研究者参考
蛋白质功能预测是药物研发中的关键任务,显著影响有效且安全药物的开发。传统机器学习模型常因蛋白质结构的复杂性和多样性而表现不佳,亟需更先进的方法。本文提出 Protein-Mamba,一种新颖的两阶段模型,结合自监督学习与微调以提升蛋白质功能预测性能。预训练阶段利用大规模无标签数据捕捉蛋白质的通用化学结构与关系;微调阶段则通过特定标注数据优化预测能力。大量实验表明,Protein-Mamba 在多个蛋白质功能数据集上的表现优于若干前沿方法。该模型有效融合无标签与有标签数据,凸显自监督学习在蛋白质功能预测中的潜力,为药物发现研究提供新方向。
原文摘要 · Abstract (English)
Protein function prediction is a pivotal task in drug discovery, significantly impacting the development of effective and safe therapeutics. Traditional machine learning models often struggle with the complexity and variability inherent in predicting protein functions, necessitating more sophisticated approaches. In this work, we introduce Protein-Mamba, a novel two-stage model that leverages both self-supervised learning and fine-tuning to improve protein function prediction. The pre-training stage allows the model to capture general chemical structures and relationships from large, unlabeled datasets, while the fine-tuning stage refines these insights using specific labeled datasets, resulting in superior prediction performance. Our extensive experiments demonstrate that Protein-Mamba achieves competitive performance, compared with a couple of state-of-the-art methods across a range of protein function datasets. This model's ability to effectively utilize both unlabeled and labeled data highlights the potential of self-supervised learning in advancing protein function prediction and offers a promising direction for future research in drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。