用可微分图划分让蛋白质语言模型的预测结果更可解释。
Structural Interpretations of Protein Language Model Representations via Differentiable Graph Partitioning

- 将语言模型特征投射到接触图,通过可微分聚类学习功能结构。
- 酶分类准确率达92.8%,结合位点检测AUC提升至0.983。
- 无需标注即可自动识别活性位点,适合生物功能研究者使用。
蛋白质语言模型如ESM-2在功能预测中表现优异,但其特征难以解释,因结构与进化信号编码于密集隐空间。本文提出即插即用框架,将ESM-2表示投影至蛋白质接触图,采用轻量级图同构网络SoftBlobGIN,通过可微分Gumbel-softmax子结构池化实现结构感知消息传递,并学习粗粒度功能子结构。在酶分类任务中,准确率92.8%,宏平均F1达0.898。与事后分析不同,该方法直接生成可审计的结构解释:GNNExplainer成功恢复生物学意义的活性位点残基、空间局域功能簇及催化接触模式。在结合位点检测中,残基AUROC从仅用ESM-2线性探针的0.885提升至0.983,表明结构解释不可由语言模型特征单独获得。学习得到的团块分区进一步提供可解释性,含已标注活性位点残基的团块重要性高出其他团块1.85倍(ρ=0.339, p=0.009),且无活性位点监督。该框架无需重训练语言模型,仅增加约110万参数,泛化性强,在ProteinShake任务上达到最大F1 0.733(GO预测)与结合位点检测AUROC 0.969。本工作定位为蛋白语言模型的可解释结构伴侣,使预测过程更透明、可审计。
原文摘要 · Abstract (English)
Protein language models such as ESM-2 learn rich residue representations that achieve strong performance on protein function prediction, but their features remain difficult to interpret as structural $\&$ evolutionary signals are encoded in dense latent spaces. We propose a plug-$\&$-play framework that projects ESM-2 representations onto protein contact graphs $\&$ applies $\textbf{SoftBlobGIN}$, a lightweight Graph Isomorphism Network with differentiable Gumbel-softmax substructure pooling, to perform structure-aware message passing $\&$ learn coarse functional substructures for downstream prediction tasks. Across enzyme classification, SoftBlobGIN achieves 92.8\% accuracy $\&$ 0.898 macro-F1. Unlike post hoc analysis of protein language models alone, our method produces directly auditable structural explanations: GNNExplainer recovers biologically meaningful active-site residues, spatially localized functional clusters, $\&$ catalytic contact patterns. On binding-site detection, SoftBlobGIN improves residue AUROC from $0.885$ using an ESM-2 linear probe to $0.983$, indicating that these structural explanations are not recoverable from language-model features alone. Learned blob partitions provide an additional layer of interpretability by automatically grouping residues into functional substructures, with blobs containing annotated active-site residues showing $1.85\times$ higher importance than other blobs ($ρ{=}0.339$, $p{=}0.009$), without any active-site supervision. Our framework requires no retraining of the language model, adds only $\sim$1.1M parameters, $\&$ generalises across ProteinShake tasks, achieving $F_{\max}$ of $0.733$ on Gene Ontology prediction $\&$ AUROC of $0.969$ on binding-site detection. We position this as an interpretable structural companion to protein language models that makes their predictions more transparent $\&$ auditable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。