研究稀疏自编码器在不同场景下能否识别可回答问题,发现效果不稳定。
Do Sparse Autoencoders Generalize? A Case Study of Answerability
- 用稀疏自编码器提取语言模型特征,测试其跨数据集泛化能力。
- 部分场景下特征表现随机,部分则优于传统探测方法。
- 适合关注模型可解释性评估与特征泛化的研究者阅读。
稀疏自编码器(SAEs)在语言模型可解释性中展现出潜力,能无监督提取稀疏特征。为使可解释性方法有效,需在不同领域识别出抽象特征,而这些特征在不同上下文中可能呈现差异。本文以“可回答性”——模型识别可回答问题的能力——为案例,对Gemma 2的SAE在多种部分自构建的可回答性数据集上进行广泛评估。分析显示,残差流探测在域内表现优于SAE特征,但跨域泛化表现差异显著:部分情况下性能接近随机,部分则超越残差流探测。结果表明,亟需更稳健的评估方法与量化工具来预测基于SAE的特征泛化能力。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) have emerged as a promising approach in language model interpretability, offering unsupervised extraction of sparse features. For interpretability methods to succeed, they must identify abstract features across domains, and these features can often manifest differently in each context. We examine this through "answerability" - a model's ability to recognize answerable questions. We extensively evaluate SAE feature generalization across diverse, partly self-constructed answerability datasets for Gemma 2 SAEs. Our analysis reveals that residual stream probes outperform SAE features within domains, but generalization performance differs sharply. SAE features show inconsistent out-of-domain transfer, with performance varying from almost random to outperforming residual stream probes. Overall, this demonstrates the need for robust evaluation methods and quantitative approaches to predict feature generalization in SAE-based interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。