发现稀疏自编码器中特征吸收现象,影响模型可解释性
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
- 通过检测特征分裂中的吸收现象,揭示稀疏自编码器的不稳定性
- 实证显示约40%的单义特征未在预期位置激活,反而被子特征吸收
- 适用于研究大模型可解释性的研究人员,尤其关注特征分解可靠性
稀疏自编码器(SAE)旨在将大语言模型(LLM)的激活空间分解为人类可理解的潜在方向或特征。随着SAE中特征数量增加,层次化特征往往分裂为更细粒度的特征(如“数学”分裂为“代数”“几何”等),这一现象称为特征分裂。然而,我们发现稀疏分解与层次特征分裂并不稳健:看似单义的特征在应触发位置未能激活,反而被其子特征“吸收”。我们提出该现象为特征吸收,并证明其根源在于优化稀疏性时,若底层特征呈层次结构,就会导致此问题。我们引入一个检测吸收的指标,并在数百个LLM的SAE上验证了这一发现。研究显示,调整SAE规模或稀疏度无法解决该问题。我们讨论了特征吸收对可解释性的影响,并提出未来需解决其根本理论缺陷,以实现大规模、可靠的模型解读。
原文摘要 · Abstract (English)
Sparse Autoencoders (SAEs) aim to decompose the activation space of large language models (LLMs) into human-interpretable latent directions or features. As we increase the number of features in the SAE, hierarchical features tend to split into finer features ("math" may split into "algebra", "geometry", etc.), a phenomenon referred to as feature splitting. However, we show that sparse decomposition and splitting of hierarchical features is not robust. Specifically, we show that seemingly monosemantic features fail to fire where they should, and instead get "absorbed" into their children features. We coin this phenomenon feature absorption, and show that it is caused by optimizing for sparsity in SAEs whenever the underlying features form a hierarchy. We introduce a metric to detect absorption in SAEs, and validate our findings empirically on hundreds of LLM SAEs. Our investigation suggests that varying SAE sizes or sparsity is insufficient to solve this issue. We discuss the implications of feature absorption in SAEs and some potential approaches to solve the fundamental theoretical issues before SAEs can be used for interpreting LLMs robustly and at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。