测试稀疏自编码器在困难任务中是否真能提升表现,结果不乐观。
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
- 用稀疏自编码器提取可解释的语义特征,尝试改进探针性能
- 在数据稀缺、标签噪声等条件下,效果不如简单基线方法稳定
- 揭示当前可解释性方法需更严格评估,适合关注模型可信度的研究者
稀疏自编码器(SAEs)是解释大语言模型(LLM)激活中概念表示的流行方法。然而,由于缺乏模型所用概念的真实标签,其解释有效性尚无充分证据,且近期研究指出了现有SAE存在的问题。一种替代验证方式是证明SAEs在下游任务上优于现有基线。本文在四种挑战性场景下测试了SAEs:数据稀缺、类别不平衡、标签噪声和协变量偏移。由于这些场景下概念检测难度高,我们假设具备可解释概念级隐变量的SAEs应提供有益归纳偏置。但实验发现,尽管SAEs在个别数据集上表现更优,却无法设计出一致超越纯基线集成的方法。此外,虽然初期看似有助于识别虚假相关、检测数据质量差及训练多标记探针,但简单非SAE基线也能达到相似效果。尽管不能排除SAEs在其他任务上的价值,本研究凸显了当前SAEs的局限性,强调必须在强基线对比下对可解释性方法进行严格评估。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) are a popular method for interpreting concepts represented in large language model (LLM) activations. However, there is a lack of evidence regarding the validity of their interpretations due to the lack of a ground truth for the concepts used by an LLM, and a growing number of works have presented problems with current SAEs. One alternative source of evidence would be demonstrating that SAEs improve performance on downstream tasks beyond existing baselines. We test this by applying SAEs to the real-world task of LLM activation probing in four regimes: data scarcity, class imbalance, label noise, and covariate shift. Due to the difficulty of detecting concepts in these challenging settings, we hypothesize that SAEs' basis of interpretable, concept-level latents should provide a useful inductive bias. However, although SAEs occasionally perform better than baselines on individual datasets, we are unable to design ensemble methods combining SAEs with baselines that consistently outperform ensemble methods solely using baselines. Additionally, although SAEs initially appear promising for identifying spurious correlations, detecting poor dataset quality, and training multi-token probes, we are able to achieve similar results with simple non-SAE baselines as well. Though we cannot discount SAEs' utility on other tasks, our findings highlight the shortcomings of current SAEs and the need to rigorously evaluate interpretability methods on downstream tasks with strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。