提出新方法评估稀疏自编码器特征的敏感性,发现许多可解释特征实际不敏感。
Measuring Sparse Autoencoder Feature Sensitivity
- 用语言模型生成语义相似文本,测试特征在相似内容上的激活稳定性
- 发现多数可解释特征对相似文本激活率低,敏感性普遍较差
- 适合关注模型可解释性与架构设计的研究者
稀疏自编码器(SAE)特征已成为机制可解释性研究的核心工具。传统方法通过分析激活样本(通常为单义且符合人类理解概念)来表征特征,但无法揭示特征敏感性:即特征在类似激活样本的文本上是否稳定激活。本文提出一种可扩展的方法评估敏感性,不依赖人工描述,而是利用语言模型生成具有相同语义属性的文本,并测试特征在这些生成文本上的激活情况。结果表明,敏感性衡量了特征质量的新维度,许多看似可解释的特征实际敏感性差。人工评估确认,当特征未能在生成文本上激活时,这些文本确实与原始激活样本高度相似。进一步研究发现,在7种不同宽度的SAE中,平均特征敏感性随模型宽度增加而下降。本工作确立敏感性作为评估个体特征和SAE架构的新标准。
原文摘要 · Abstract (English)
Sparse Autoencoder (SAE) features have become essential tools for mechanistic interpretability research. SAE features are typically characterized by examining their activating examples, which are often "monosemantic" and align with human interpretable concepts. However, these examples don't reveal feature sensitivity: how reliably a feature activates on texts similar to its activating examples. In this work, we develop a scalable method to evaluate feature sensitivity. Our approach avoids the need to generate natural language descriptions for features; instead we use language models to generate text with the same semantic properties as a feature's activating examples. We then test whether the feature activates on these generated texts. We demonstrate that sensitivity measures a new facet of feature quality and find that many interpretable features have poor sensitivity. Human evaluation confirms that when features fail to activate on our generated text, that text genuinely resembles the original activating examples. Lastly, we study feature sensitivity at the SAE level and observe that average feature sensitivity declines with increasing SAE width across 7 SAE variants. Our work establishes feature sensitivity as a new dimension for evaluating both individual features and SAE architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。