用稀疏自编码器解析抗体语言模型的内部机制,实现精准生成控制。
Mechanistic Interpretability of Antibody Language Models Using SAEs
- 采用TopK与有序稀疏自编码器分析抗体模型的潜在特征。
- 有序SAE能可靠识别可操控的生成特征,但激活模式更复杂。
- 适合需要精确生成引导的研究者,如抗体设计与功能调控。
稀疏自编码器(SAEs)是一种机制可解释性技术,可用于揭示大型蛋白质语言模型中学习到的概念。本文使用TopK和有序稀疏自编码器研究自回归抗体语言模型,并探索其生成过程的可控性。结果表明,TopK SAE可揭示生物学上有意义的潜在特征,但高特征-概念相关性并不保证对生成过程的因果控制。相比之下,有序SAE引入层次结构,能可靠识别可操控特征,但导致激活模式更复杂且不易解释。这些发现推进了领域特定蛋白质语言模型的机制可解释性研究,表明虽然TopK SAE足以将潜在特征映射到概念,但在需要精确生成引导时,有序SAE更为合适。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) are a mechanistic interpretability technique that have been used to provide insight into learned concepts within large protein language models. Here, we employ TopK and Ordered SAEs to investigate autoregressive antibody language models, and steer their generation. We show that TopK SAEs can reveal biologically meaningful latent features, but high feature-concept correlation does not guarantee causal control over generation. In contrast, Ordered SAEs impose a hierarchical structure that reliably identifies steerable features, but at the expense of more complex and less interpretable activation patterns. These findings advance the mechanistic interpretability of domain-specific protein language models and suggest that, while TopK SAEs suffice for mapping latent features to concepts, Ordered SAEs are preferable when precise generative steering is required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。