arXiv:2504.07156q-bio.BMcs.AI2025-04被引 5

将蛋白语言模型嵌入分解为可解释与保留性能的两部分。

PLM-eXplain: Divide and Conquer the Protein Embedding Space

  • 将PLM嵌入拆分为生物化学特征可解释子空间和剩余子空间。
  • 在三个蛋白分类任务中保持高精度,同时实现决策可解释。
  • 适合需要理解模型推理过程的生物医学研究者使用。

蛋白质语言模型(PLMs)通过生成强大的序列表示,革新了计算生物学。然而其黑箱特性限制了生物学解释和可行动洞察的转化。我们提出可解释适配层PLM-eXplain(PLM-X),将PLM嵌入分解为基于已知生化特征的可解释子空间和保留模型预测能力的残差子空间。基于ESM2嵌入,该适配器整合了二级结构、疏水性等经典属性,同时维持高性能。我们在三个蛋白级分类任务中验证了有效性:外泌体关联预测、跨膜螺旋识别和聚集倾向预测。PLM-X在不牺牲准确率的前提下实现了模型决策的生物学可解释性,为多种下游应用提供了通用的可解释性增强方案。本工作填补了强大深度学习模型与可行动生物学见解之间的关键空白。

原文摘要 · Abstract (English)

Protein language models (PLMs) have revolutionised computational biology through their ability to generate powerful sequence representations for diverse prediction tasks. However, their black-box nature limits biological interpretation and translation to actionable insights. We present an explainable adapter layer - PLM-eXplain (PLM-X), that bridges this gap by factoring PLM embeddings into two components: an interpretable subspace based on established biochemical features, and a residual subspace that preserves the model's predictive power. Using embeddings from ESM2, our adapter incorporates well-established properties, including secondary structure and hydropathy while maintaining high performance. We demonstrate the effectiveness of our approach across three protein-level classification tasks: prediction of extracellular vesicle association, identification of transmembrane helices, and prediction of aggregation propensity. PLM-X enables biological interpretation of model decisions without sacrificing accuracy, offering a generalisable solution for enhancing PLM interpretability across various downstream applications. This work addresses a critical need in computational biology by providing a bridge between powerful deep learning models and actionable biological insights.

蛋白语言模型可解释性生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。