对比通用与领域模型,发现结构信息比预训练更重要
General Protein Pretraining or Domain-Specific Designs? Benchmarking Protein Modeling on Realistic Applications
- 构建跨五类任务的蛋白模型评估基准Protap
- 小数据监督训练常优于大规模预训练
- 结构信息与生物先验能提升特定任务性能
近期大量深度学习架构和预训练策略被用于支持下游蛋白应用。同时,融合生物知识的领域专用模型也得到发展。本文提出Protap,一个系统比较主干架构、预训练策略及领域专用模型的综合性基准,涵盖三类通用任务和两类新设工业相关任务:酶催化蛋白切割位点预测与靶向蛋白降解。每个任务均在多种预训练设置下对比不同领域模型与通用架构。实证研究表明:(i) 尽管大规模预训练编码器表现优异,但在小规模下游数据集上监督训练的编码器常表现更优;(ii) 在下游微调中引入结构信息可达到甚至超过在大规模序列语料上预训练的蛋白语言模型;(iii) 融入领域生物先验能显著提升特定任务性能。代码与数据集已公开于https://github.com/Trust-App-AI-Lab/protap。
原文摘要 · Abstract (English)
Recently, extensive deep learning architectures and pretraining strategies have been explored to support downstream protein applications. Additionally, domain-specific models incorporating biological knowledge have been developed to enhance performance in specialized tasks. In this work, we introduce $\textbf{Protap}$, a comprehensive benchmark that systematically compares backbone architectures, pretraining strategies, and domain-specific models across diverse and realistic downstream protein applications. Specifically, Protap covers five applications: three general tasks and two novel specialized tasks, i.e., enzyme-catalyzed protein cleavage site prediction and targeted protein degradation, which are industrially relevant yet missing from existing benchmarks. For each application, Protap compares various domain-specific models and general architectures under multiple pretraining settings. Our empirical studies imply that: (i) Though large-scale pretraining encoders achieve great results, they often underperform supervised encoders trained on small downstream training sets. (ii) Incorporating structural information during downstream fine-tuning can match or even outperform protein language models pretrained on large-scale sequence corpora. (iii) Domain-specific biological priors can enhance performance on specialized downstream tasks. Code and datasets are publicly available at https://github.com/Trust-App-AI-Lab/protap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。