arXiv:2509.05309q-bio.QMcs.AI2025-09AAAI被引 3

用语义引导提升蛋白语言模型的可解释性

ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders

  • 训练时结合注释数据与领域知识,解耦神经元语义混杂
  • 相比旧方法,特征更生物相关且可解释性更强
  • 适合想理解蛋白模型内部机制的研究者

稀疏自编码器(SAE)已成为大语言模型可解释性分析的重要工具。近期研究将SAE应用于蛋白语言模型(PLMs),试图从其隐空间提取并分析具有生物学意义的特征。然而,现有SAE存在语义混杂问题,单个神经元常同时表征多个非线性概念,导致模型行为难以可靠解读或操控。本文提出一种语义引导的稀疏自编码器(ProtSAE)。不同于以往依赖标注数据筛选和解释激活的方法,我们通过结合标注数据与领域知识,在训练阶段即引导语义解耦,以缓解属性混杂的影响。通过可解释性实验表明,ProtSAE学习到的隐层特征比先前方法更具生物学相关性和可解释性。性能分析进一步显示,ProtSAE在保持高重建保真度的同时,在可解释探测任务上表现更优。此外,我们还展示了ProtSAE在下游生成任务中引导PLMs的潜力。

原文摘要 · Abstract (English)

Sparse Autoencoder (SAE) has emerged as a powerful tool for mechanistic interpretability of large language models. Recent works apply SAE to protein language models (PLMs), aiming to extract and analyze biologically meaningful features from their latent spaces. However, SAE suffers from semantic entanglement, where individual neurons often mix multiple nonlinear concepts, making it difficult to reliably interpret or manipulate model behaviors. In this paper, we propose a semantically-guided SAE, called ProtSAE. Unlike existing SAE which requires annotation datasets to filter and interpret activations, we guide semantic disentanglement during training using both annotation datasets and domain knowledge to mitigate the effects of entangled attributes. We design interpretability experiments showing that ProtSAE learns more biologically relevant and interpretable hidden features compared to previous methods. Performance analyses further demonstrate that ProtSAE maintains high reconstruction fidelity while achieving better results in interpretable probing. We also show the potential of ProtSAE in steering PLMs for downstream generation tasks.

可解释性蛋白语言模型稀疏自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。