用稀疏自编码器解析蛋白质模型内部机制并引导序列生成
Interpreting and Steering Protein Language Models through Sparse Autoencoders
- 通过稀疏自编码器分析蛋白语言模型的潜在表征
- 发现特定潜变量与跨膜区、结合位点等特征相关
- 可筛选潜变量实现对锌指结构等目标的定向生成
基于Transformer的自然语言模型发展迅速,但其内部机制仍难理解。本文将稀疏自编码器(SAE)应用于800万参数的ESM-2蛋白语言模型,通过统计每个潜在成分对蛋白质注释的相关性,识别出与跨膜区域、结合位点及特殊基序相关的潜在表征。利用这些发现,可筛选出与目标特征相关的潜变量,引导模型生成特定结构,如锌指域。该研究为生物序列模型的机制可解释性提供了新视角,推动了序列设计中的模型操控能力。
原文摘要 · Abstract (English)
The rapid advancements in transformer-based language models have revolutionized natural language processing, yet understanding the internal mechanisms of these models remains a significant challenge. This paper explores the application of sparse autoencoders (SAE) to interpret the internal representations of protein language models, specifically focusing on the ESM-2 8M parameter model. By performing a statistical analysis on each latent component's relevance to distinct protein annotations, we identify potential interpretations linked to various protein characteristics, including transmembrane regions, binding sites, and specialized motifs. We then leverage these insights to guide sequence generation, shortlisting the relevant latent components that can steer the model towards desired targets such as zinc finger domains. This work contributes to the emerging field of mechanistic interpretability in biological sequence models, offering new perspectives on model steering for sequence design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。