通过共进化信息增强蛋白序列表示,提升模型泛化能力。
SFM-Protein: Integrative Co-evolutionary Pre-training for Advanced Protein Sequence Representation
- 基于残基间相互作用设计共进化预训练策略
- 在多个下游任务中超越同类模型表现
- 适合蛋白质结构与功能研究者使用
蛋白质在生物系统中至关重要,其功能与其三维结构密切相关。理解蛋白质结构与氨基酸序列之间的关系仍是蛋白建模的核心挑战。传统蛋白基础模型虽通过大规模无标注数据预训练获益,但常难以捕捉关键的共进化信息,而基于进化的方法在此方面表现更优。本研究提出一种新型蛋白基础模型预训练策略,聚焦氨基酸残基间的相互作用,以增强从序列数据中提取短程与长程共进化特征的能力。模型在大规模蛋白序列数据集上训练,展现出优异的泛化性能,在多种下游任务中优于同等规模的现有基线模型(包括ESM),实验证明其有效整合了共进化信息,显著推进了基于序列的蛋白建模发展。
原文摘要 · Abstract (English)
Proteins, essential to biological systems, perform functions intricately linked to their three-dimensional structures. Understanding the relationship between protein structures and their amino acid sequences remains a core challenge in protein modeling. While traditional protein foundation models benefit from pre-training on vast unlabeled datasets, they often struggle to capture critical co-evolutionary information, which evolutionary-based methods excel at. In this study, we introduce a novel pre-training strategy for protein foundation models that emphasizes the interactions among amino acid residues to enhance the extraction of both short-range and long-range co-evolutionary features from sequence data. Trained on a large-scale protein sequence dataset, our model demonstrates superior generalization ability, outperforming established baselines of similar size, including the ESM model, across diverse downstream tasks. Experimental results confirm the model's effectiveness in integrating co-evolutionary information, marking a significant step forward in protein sequence-based modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。