arXiv:2505.16896cs.LGcs.AI2025-05被引 1

用图神经网络增强语言模型的结构感知能力,提升蛋白质功能预测效果。

Structure-Aligned Protein Language Model

  • 通过对比学习对齐语言模型与图神经网络的残基表征,注入跨蛋白结构信息。
  • 联合训练预测结构标记,融合蛋白内部结构特征,在CASP16上接触预测提升59%。
  • 方法轻量易用,适配大模型,对突变效应预测等任务均有显著增益。

基于大规模蛋白序列数据库预训练的蛋白质语言模型(pLMs)在多种下游任务中表现优异,但在某些生物应用中缺乏必要的结构知识。为此,我们提出一种方法,利用预训练的蛋白质图神经网络(pGNNs)为pLMs注入结构知识。首先,通过跨多个蛋白质的残基表示对比学习,实现语言模型与图神经网络的潜层对齐,引入跨蛋白结构信息;其次,通过物理层面的任务训练语言模型预测结构标记,整合蛋白内部结构信息。该双任务框架有效融合了跨蛋白与蛋白内结构知识。针对PDB中结构质量参差不齐的问题,进一步设计残基损失选择模块,使用小模型筛选高质量且具有挑战性的残基损失供语言模型学习。将该结构对齐方法作为轻量级后训练步骤应用于当前最先进的ESM2和AMPLIFY模型,显著提升性能:在深突变扫描(DMS)拟合度预测中取得显著进步,且在CASP16上,ESM2 650M的接触预测精确率(P@L)提升59%。这些改进在8M至650M模型规模范围内均保持稳定,并扩展至多种下游任务。

原文摘要 · Abstract (English)

Protein language models (pLMs) pre-trained on vast protein sequence databases excel at various downstream tasks but often lack the structural knowledge essential for some biological applications. To address this, we introduce a method to enrich pLMs with structural knowledge by leveraging pre-trained protein graph neural networks (pGNNs). First, a latent-level contrastive learning task aligns residue representations from pLMs with those from pGNNs across multiple proteins, injecting inter-protein structural information. Additionally, a physical-level task integrates intra-protein information by training pLMs to predict structure tokens. Together, the proposed dual-task framework effectively incorporates both inter- and intra-protein structural knowledge into pLMs. Given the variability in the quality of protein structures in PDB, we further introduce a residue loss selection module that uses a small model trained on high-quality structures to select reliable yet challenging residue losses for the pLM to learn. Applying our structure alignment method as a simple, lightweight post-training step to the state-of-the-art ESM2 and AMPLIFY yields notable performance gains. These improvements are consistent across a wide range of tasks, including substantial gains in deep mutational scanning (DMS) fitness prediction and a 59% increase in P@L for ESM2 650M contact prediction on CASP16. Furthermore, we demonstrate that these performance gains are robust, scaling with model sizes from 8M to 650M and extending to different downstream tasks.

蛋白质语言模型图神经网络结构预测多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。