arXiv:2608.25548cs.LGq-bio.QM2026-08

用正交投影分离蛋白语言模型中的生化特征,揭示其对蛋白适应度预测的贡献。

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

论文配图:Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction
图 1 · 摘自论文原文
  • 通过正交投影去除嵌入中已知生化特征的线性及高阶影响。
  • 移除特征后分类器性能下降,说明嵌入编码了相关生化模式。
  • 方法高效可迁移,适用于其他生物序列预测任务。

近年来,蛋白语言模型(PLMs)在生物医学领域得到广泛应用。其嵌入为蛋白序列提供了丰富的数值表示,在蛋白适应度预测等下游任务中表现卓越。然而,PLM嵌入缺乏直接可解释性,难以明确其编码的具体生化特征。为此,我们采用正交投影技术,移除嵌入中已知表格特征的线性及高阶交互效应。通过消融实验发现,仅使用去除非可解释部分后的嵌入训练分类器时,性能显著下降。进一步评估表明,这些生化特征解释了分类器预测结果中相当大一部分方差。因此,我们证明了PLM嵌入编码了与生化特性相关的模式,并量化了其对蛋白适应度预测的贡献。该计算高效的框架不限于当前特征或嵌入,可轻松推广至其他生物序列预测场景。

原文摘要 · Abstract (English)

Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on several downstream tasks including protein fitness prediction. However, PLM embeddings are not directly interpretable and, thereby, it remains unclear what features they encode. To gain insight into which biochemical properties of the protein are driving the prediction, we leverage an orthogonal projection technique that removes linear effects of known tabular features from embeddings and extend it to high-order and interaction effects. In this way, we remove the effects of interpretable biochemical features from PLM embeddings. In an ablation study, we show that this leads to a decrease in performance for a downstream classifier trained only on the embeddings to predict protein fitness. In an additional evaluation, we find that these biochemical features explain a substantial part of the variance in the predictions of this classifier. Hence, we can show that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness. This computationally efficient approach is not limited to the features or embeddings considered here and is readily transferable to problem settings beyond protein fitness prediction.

蛋白语言模型可解释性嵌入分析适应度预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。