arXiv:2410.22949cs.LGq-bio.BM2024-10NeurIPS被引 4

MutaPLM让蛋白突变预测可解释,助力精准蛋白设计。

MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering

  • 构建蛋白差分网络,显式捕捉突变特征
  • 通过文本链式推理挖掘突变知识,提升解释性
  • 首个带丰富文本标注的大规模突变数据集,适合生物研究者

研究蛋白质序列中的氨基酸突变在生命科学中具有重要意义。蛋白语言模型(PLMs)虽在诸多生物学应用中表现优异,但因架构设计与缺乏监督,通常仅隐式建模突变,依赖进化合理性,难以作为可解释、可工程化的工具。为此,我们提出MutaPLM,一个统一的蛋白突变解释与导航框架。该框架引入蛋白差分网络,在统一特征空间中显式表示突变,并采用链式思维(CoT)策略的迁移学习流程,从生物医学文本中提取突变知识。我们还构建了MutaDescribe,首个大规模蛋白突变数据集,包含丰富的文本注释,提供跨模态监督信号。实验表明,MutaPLM能有效生成人类可理解的突变影响解释,并优先筛选出具有理想性质的新突变。代码、模型与数据已开源。

原文摘要 · Abstract (English)

Studying protein mutations within amino acid sequences holds tremendous significance in life sciences. Protein language models (PLMs) have demonstrated strong capabilities in broad biological applications. However, due to architectural design and lack of supervision, PLMs model mutations implicitly with evolutionary plausibility, which is not satisfactory to serve as explainable and engineerable tools in real-world studies. To address these issues, we present MutaPLM, a unified framework for interpreting and navigating protein mutations with protein language models. MutaPLM introduces a protein delta network that captures explicit protein mutation representations within a unified feature space, and a transfer learning pipeline with a chain-of-thought (CoT) strategy to harvest protein mutation knowledge from biomedical texts. We also construct MutaDescribe, the first large-scale protein mutation dataset with rich textual annotations, which provides cross-modal supervision signals. Through comprehensive experiments, we demonstrate that MutaPLM excels at providing human-understandable explanations for mutational effects and prioritizing novel mutations with desirable properties. Our code, model, and data are open-sourced at https://github.com/PharMolix/MutaPLM.

蛋白语言模型突变解释可解释AI生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。