用3D残基图与关系感知注意力,提升蛋白质二级结构预测精度。
Protein Secondary Structure Prediction Using 3D Graphs and Relation-Aware Message Passing Transformers
- 构建蛋白残基图,融合序列与三维结构关系进行消息传递
- 在3种和8种二级结构分类上均超越基线模型的F1分数
- 适合需要高精度结构预测的生物信息学研究者使用
本研究针对从蛋白质一级序列预测二级结构这一关键任务,旨在为三级结构预测提供基础,并揭示蛋白功能与相互作用。现有方法多依赖大量未标注序列,却未显式利用对功能起决定性作用的3D结构数据。为此,我们引入蛋白残基图,设计多种序列或结构连接方式以捕捉更丰富的空间信息。结合预训练的Transformer型蛋白语言模型编码序列,采用GCN与R-GCN等消息传递机制学习几何特征。通过在特定节点邻域内执行卷积并堆叠多层,有效融合空间图中的复杂关联与依赖关系。在NetSurfP-2.0提供的训练数据集上评估,该模型在3种和8种二级结构分类下均取得优于基线的F1分数。
原文摘要 · Abstract (English)
In this study, we tackle the challenging task of predicting secondary structures from protein primary sequences, a pivotal initial stride towards predicting tertiary structures, while yielding crucial insights into protein activity, relationships, and functions. Existing methods often utilize extensive sets of unlabeled amino acid sequences. However, these approaches neither explicitly capture nor harness the accessible protein 3D structural data, which is recognized as a decisive factor in dictating protein functions. To address this, we utilize protein residue graphs and introduce various forms of sequential or structural connections to capture enhanced spatial information. We adeptly combine Graph Neural Networks (GNNs) and Language Models (LMs), specifically utilizing a pre-trained transformer-based protein language model to encode amino acid sequences and employing message-passing mechanisms like GCN and R-GCN to capture geometric characteristics of protein structures. Employing convolution within a specific node's nearby region, including relations, we stack multiple convolutional layers to efficiently learn combined insights from the protein's spatial graph, revealing intricate interconnections and dependencies in its structural arrangement. To assess our model's performance, we employed the training dataset provided by NetSurfP-2.0, which outlines secondary structure in 3-and 8-states. Extensive experiments show that our proposed model, SSRGNet surpasses the baseline on f1-scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。