用稀疏自编码器挖掘蛋白质语言模型的可解释特征。
InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
- 通过稀疏自编码器从ESM-2嵌入中提取可读特征。
- 每层发现最多2548个特征,关联143个生物学概念。
- 适合生物信息学、蛋白设计与模型可解释性研究者。
蛋白质语言模型(PLMs)在蛋白质建模与设计中表现卓越,但其预测结构与功能的内部机制仍不清晰。本文提出系统方法,利用稀疏自编码器(SAEs)从PLM ESM-2的嵌入中提取并分析可解释特征。训练SAEs后,每层识别出最多2,548个与143个已知生物学概念(如结合位点、结构基序、功能域)强相关的可解释潜变量。相比之下,单独分析ESM-2神经元仅发现每层最多46个具明确概念对齐的神经元,表明大多数概念以叠加方式表示。除覆盖已有注释外,还发现未映射到现有注释的连贯新概念,并提出基于语言模型自动解读新潜变量的流程。实际应用中,这些特征可用于填补蛋白数据库缺失注释,并实现对蛋白质序列生成的定向调控。结果表明PLMs编码了丰富且可解释的蛋白质生物学表征,本文构建了系统框架用于提取与分析。作为社区资源,发布交互式可视化平台InterPLM(interPLM.ai)及训练分析代码(github.com/ElanaPearl/interPLM)。
原文摘要 · Abstract (English)
Protein language models (PLMs) have demonstrated remarkable success in protein modeling and design, yet their internal mechanisms for predicting structure and function remain poorly understood. Here we present a systematic approach to extract and analyze interpretable features from PLMs using sparse autoencoders (SAEs). By training SAEs on embeddings from the PLM ESM-2, we identify up to 2,548 human-interpretable latent features per layer that strongly correlate with up to 143 known biological concepts such as binding sites, structural motifs, and functional domains. In contrast, examining individual neurons in ESM-2 reveals up to 46 neurons per layer with clear conceptual alignment across 15 known concepts, suggesting that PLMs represent most concepts in superposition. Beyond capturing known annotations, we show that ESM-2 learns coherent concepts that do not map onto existing annotations and propose a pipeline using language models to automatically interpret novel latent features learned by the SAEs. As practical applications, we demonstrate how these latent features can fill in missing annotations in protein databases and enable targeted steering of protein sequence generation. Our results demonstrate that PLMs encode rich, interpretable representations of protein biology and we propose a systematic framework to extract and analyze these latent features. In the process, we recover both known biology and potentially new protein motifs. As community resources, we introduce InterPLM (interPLM.ai), an interactive visualization platform for exploring and analyzing learned PLM features, and release code for training and analysis at github.com/ElanaPearl/interPLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。