融合蛋白序列与结构信息,提升多模态蛋白表征性能
Bidirectional Hierarchical Protein Multi-Modal Representation Learning
- 双向分层融合机制,打通序列与结构模态的信息通道
- 在酶分类、结合亲和力等6项任务上超越现有方法
- 适合需要综合序列与结构信息的研究者使用
蛋白质表征学习对众多生物任务至关重要。近期基于大规模蛋白质序列预训练的Transformer模型(pLMs)在序列相关任务中表现优异,但缺乏结构上下文;而图神经网络(GNNs)虽能利用三维结构信息并在预测任务中展现良好泛化能力,却受限于标注结构数据稀缺。鉴于序列与结构是同一蛋白实体的互补视角,我们提出一种多模态双向分层融合框架,通过注意力与门控机制实现pLMs生成的序列表征与GNN提取的结构特征之间的有效交互,增强神经网络各层的信息交换与融合。该双向分层(Bi-Hierarchical)融合策略充分发挥双模态优势,捕获更丰富全面的蛋白质表征。在此基础上,我们设计局部门控融合与全局多头自注意力融合两种方法。实验表明,本方法在酶EC分类、模型质量评估、蛋白-配体结合亲和力预测、蛋白-蛋白结合位点预测及B细胞表位预测等多项基准任务中持续优于强基线与现有融合技术,建立了多模态蛋白质表征学习的新基准,验证了双向分层融合在连接序列与结构模态方面的有效性。
原文摘要 · Abstract (English)
Protein representation learning is critical for numerous biological tasks. Recently, large transformer-based protein language models (pLMs) pretrained on large scale protein sequences have demonstrated significant success in sequence-based tasks. However, pLMs lack structural context. Conversely, graph neural networks (GNNs) designed to leverage 3D structural information have shown promising generalization in protein-related prediction tasks, but their effectiveness is often constrained by the scarcity of labeled structural data. Recognizing that sequence and structural representations are complementary perspectives of the same protein entity, we propose a multimodal bidirectional hierarchical fusion framework to effectively merge these modalities. Our framework employs attention and gating mechanisms to enable effective interaction between pLMs-generated sequential representations and GNN-extracted structural features, improving information exchange and enhancement across layers of the neural network. This bidirectional and hierarchical (Bi-Hierarchical) fusion approach leverages the strengths of both modalities to capture richer and more comprehensive protein representations. Based on the framework, we further introduce local Bi-Hierarchical Fusion with gating and global Bi-Hierarchical Fusion with multihead self-attention approaches. Our method demonstrates consistent improvements over strong baselines and existing fusion techniques in a variety of protein representation learning benchmarks, including enzyme EC classification, model quality assessment, protein-ligand binding affinity prediction, protein-protein binding site prediction, and B cell epitopes prediction. Our method establishes a new state-of-the-art for multimodal protein representation learning, emphasizing the efficacy of Bi-Hierarchical Fusion in bridging sequence and structural modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。