通过多任务训练提升蛋白质语言模型的表示能力
Ankh3: Multi-Task Pretraining with Sequence Denoising and Completion Enhances Protein Representations
- 联合优化掩码建模与序列补全两个任务
- 在二级结构、荧光、适应性等任务上表现更优
- 适合需要精准蛋白质属性预测的研究者
蛋白质语言模型(PLMs)已成为识别蛋白质序列复杂模式的强大工具。然而,仅依赖单一预训练任务可能限制其捕捉序列全部信息的能力。尽管引入多模态数据或监督目标可提升性能,但预训练仍常聚焦于修复损坏序列。为突破这一局限,我们提出Ankh3,通过联合优化两种目标:使用多种掩码概率的掩码语言建模,以及仅以蛋白质序列为输入的序列补全任务。该多任务预训练策略证明,仅基于蛋白质序列即可学习到更丰富、更具泛化性的表示。实验显示,该模型在下游任务中表现显著提升,包括二级结构预测、荧光强度预测、GB1适应性及接触图预测。多任务整合使模型对蛋白质性质有更全面理解,从而实现更鲁棒、更准确的预测。
原文摘要 · Abstract (English)
Protein language models (PLMs) have emerged as powerful tools to detect complex patterns of protein sequences. However, the capability of PLMs to fully capture information on protein sequences might be limited by focusing on single pre-training tasks. Although adding data modalities or supervised objectives can improve the performance of PLMs, pre-training often remains focused on denoising corrupted sequences. To push the boundaries of PLMs, our research investigated a multi-task pre-training strategy. We developed Ankh3, a model jointly optimized on two objectives: masked language modeling with multiple masking probabilities and protein sequence completion relying only on protein sequences as input. This multi-task pre-training demonstrated that PLMs can learn richer and more generalizable representations solely from protein sequences. The results demonstrated improved performance in downstream tasks, such as secondary structure prediction, fluorescence, GB1 fitness, and contact prediction. The integration of multiple tasks gave the model a more comprehensive understanding of protein properties, leading to more robust and accurate predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。