用二值化稀疏编码提升神经网络特征可解释性
Binary Sparse Coding for Interpretability
- 将激活值强制为0或1,消除高低激活强度差异
- 二值化后特征更单一语义,但重建误差上升
- 适合关注特征语义纯粹性的可解释性研究
稀疏自编码器(SAEs)用于将神经网络激活分解为稀疏激活的特征,但许多SAE特征仅在高激活强度下可解释。为此,我们提出二值稀疏自编码器(BAEs)和二值转换器(BTCs),强制所有激活值为0或1。实验发现,二值化显著提升了特征的可解释性和单义性,同时增加了重构误差。通过消除高低激活强度的区分,避免了连续激活中隐含的不可解释信息。然而,二值化也导致大量超高频不可解释特征出现;当解释度评分考虑频率影响后,连续稀疏编码器的得分略优于二值模型。这表明多义性可能是神经激活中无法消除的特性。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) are used to decompose neural network activations into sparsely activating features, but many SAE features are only interpretable at high activation strengths. To address this issue we propose to use binary sparse autoencoders (BAEs) and binary transcoders (BTCs), which constrain all activations to be zero or one. We find that binarisation significantly improves the interpretability and monosemanticity of the discovered features, while increasing reconstruction error. By eliminating the distinction between high and low activation strengths, we prevent uninterpretable information from being smuggled in through the continuous variation in feature activations. However, we also find that binarisation increases the number of uninterpretable ultra-high frequency features, and when interpretability scores are frequency-adjusted, the scores for continuous sparse coders are slightly better than those of binary ones. This suggests that polysemanticity may be an ineliminable property of neural activations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。