arXiv:2410.02023cs.LGcs.AI2024-10中稿 · Bioinformatics被引 9

DeepProtein提供蛋白序列学习的深度学习工具与基准测试。

DeepProtein: Deep Learning Library and Benchmark for Protein Sequence Learning

  • 基于Prot-T5构建可直接使用的蛋白模型库。
  • 在4个任务上达到顶尖性能,6个任务表现优异。
  • 适合生物信息学与药物发现研究者快速上手使用。

深度学习深刻影响了蛋白质科学,推动了蛋白质属性、高级结构和分子互作预测的突破。本文介绍DeepProtein,一个专为蛋白相关任务设计的综合性、易用的深度学习库,支持研究人员无缝处理蛋白数据并应用前沿模型。为评估模型性能,我们建立了基准测试,涵盖蛋白功能预测、亚细胞定位预测、蛋白-蛋白互作预测和蛋白结构预测等多任务。此外,我们提出了DeepProt-T5系列微调模型,在四个基准任务上达到当前最优表现,并在另外六个任务上表现出色。完整文档与教程确保易用性与可复现性。DeepProtein基于广泛使用的药物发现库DeepPurpose,可在https://github.com/jiaqingxie/DeepProtein公开获取。

原文摘要 · Abstract (English)

Deep learning has deeply influenced protein science, enabling breakthroughs in predicting protein properties, higher-order structures, and molecular interactions. This paper introduces DeepProtein, a comprehensive and user-friendly deep learning library tailored for protein-related tasks. It enables researchers to seamlessly address protein data with cutting-edge deep learning models. To assess model performance, we establish a benchmark evaluating different deep learning architectures across multiple protein-related tasks, including protein function prediction, subcellular localization prediction, protein-protein interaction prediction, and protein structure prediction. Furthermore, we introduce DeepProt-T5, a series of fine-tuned Prot-T5-based models that achieve state-of-the-art performance on four benchmark tasks, while demonstrating competitive results on six of others. Comprehensive documentation and tutorials are available which could ensure accessibility and support reproducibility. Built upon the widely used drug discovery library DeepPurpose, DeepProtein is publicly available at https://github.com/jiaqingxie/DeepProtein.

蛋白学习深度学习模型库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。