用功能信息增强蛋白多模态预训练,显著提升下游任务表现。
ProtCLIP: Function-Informed Protein Multi-Modal Learning
- 基于属性采样构建高质量蛋白-文本数据集ProtAnno,平衡数据量与质量。
- 设计分段预训练目标,显式建模蛋白静态与动态功能区域。
- 在22个基准上达到最优,跨模态转换任务平均提升75%。
多模态预训练通过对齐蛋白序列与生物描述,已学习到通用的蛋白表示并在多种下游任务中表现优异。然而,由于对齐的蛋白-文本数据利用效率低,且缺乏有效的功能导向预训练范式,这类方法未能复制语言-视觉基础模型的卓越成就。为此,本文构建了大规模蛋白-文本配对数据集ProtAnno,采用基于属性的采样策略,根据样本置信度与属性覆盖率决定选择概率,有效平衡了大规模噪声数据下的数据质量与数量。此外,受蛋白特异性功能机制重要性的启发,提出的范式通过两个分段预训练目标,显式建模蛋白的静态与动态功能片段,以功能信息注入细粒度特征。结合上述创新,我们开发了ProtCLIP,一种全面表征功能感知蛋白嵌入的多模态基础模型。在涵盖5类任务(功能分类、突变效应预测、跨模态转换、语义相似性推断、蛋白质相互作用预测)的22个不同蛋白基准上,ProtCLIP持续取得最先进性能,其中跨模态转换任务平均提升75%,GO-CC任务提升59.9%,GO-BP任务提升39.7%。实验结果验证了ProtCLIP作为蛋白多模态基础模型的巨大潜力。
原文摘要 · Abstract (English)
Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。