arXiv:2411.08909q-bio.BMcs.LG2024-11被引 9

用高效架构提升蛋白模型长序列理解能力,支持复杂相互作用建模。

Long-context Protein Language Modeling Using Bidirectional Mamba with Shared Projection Layers

  • 采用双向Mamba结构实现长序列蛋白建模,共享投影层降低计算开销。
  • 在1000亿和1万亿参数下,下游任务性能较ESM-2提升30%和16%。
  • 可融入蛋白质互作图谱,适用于结构与功能预测等生物医学场景。

自监督语言模型在蛋白序列表征学习和药物生成设计中取得显著进展。现有蛋白语言模型多基于Transformer架构,仅在短序列上训练,难以外推至长蛋白或蛋白复合物,且无法充分捕捉生物分子相互作用与动态机制。本文提出基于选择性结构状态空间模型的双向Mamba-S架构,构建长上下文蛋白语言模型LC-PLM,通过掩码语言建模在氨基酸级别学习高质量通用表征。进一步引入图上下文变体LC-PLM-G,利用蛋白质互作(PPI)图进行二次训练。实验表明,LC-PLM具备良好的神经缩放规律,具有更强的长度外推能力,在使用100B和1T tokens训练时,相比基于Transformer的ESM-2,在下游任务上分别提升30%和16%。结合PPI图的LC-PLM-G在蛋白结构与功能预测任务中表现优异。研究证明,通过计算高效的结构化状态空间模型扩大上下文规模,并融合生物图谱中的分子交互信息,有助于学习更通用的蛋白表征。

原文摘要 · Abstract (English)

Self-supervised training of language models (LMs) has seen great success for protein sequences in learning meaningful representations and for generative drug design. Most protein LMs are based on the Transformer architecture trained on individual proteins with short context lengths. Such protein LMs cannot extrapolate to longer proteins and protein complexes well. They also fail to account for the underlying biological mechanisms carried out by biomolecular interactions and dynamics i.e., proteins often interact with other proteins, molecules, and pathways in complex biological systems. In this work, we propose LC-PLM based on an alternative protein LM architecture, BiMamba-S, built upon selective structured state-space models, to learn high-quality universal protein representations at the amino acid token level using masked language modeling. We also introduce its graph-contextual variant, LC-PLM, which contextualizes protein-protein interaction (PPI) graphs for a second stage of training. LC-PLM demonstrates favorable neural scaling laws, better length extrapolation capability, and up to 30% and 16% improvements on protein downstream tasks compared to Transformer-based ESM-2 when trained with 100B and 1T tokens, respectively. LC-PLM-G further trained within the context of PPI graphs shows promising results on protein structure and function prediction tasks. Our study demonstrates the benefit of increasing the context size with computationally efficient LM architecture (e.g., structured state space models) in learning universal protein representations and incorporating molecular interaction contexts contained in biological graphs.

蛋白语言模型长序列建模Mamba生物图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。