arXiv:2602.14111cs.LG2026-02被引 8

稀疏自编码器解释神经网络的能力存疑,实测其效果不如随机基线。

Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?

  • 在已知真特征的合成数据上,只恢复9%真实特征,尽管重构率71%
  • 真实激活测试中,随机基线与训练好的SAE在可解释性、探针和因果编辑上表现相当
  • 当前SAE可能无法可靠分解模型内部机制,适合研究解释性方法的可信度

稀疏自编码器(SAEs)被视为解释神经网络的有效工具,能将激活分解为稀疏且人类可读的特征。尽管近年有多种SAE变体并成功扩展至前沿模型,但下游任务中的负面结果引发了对其是否捕捉到有意义特征的质疑。为此,我们进行了两项互补评估:在具有已知真特征的合成设置中,发现即使重建方差达71%,SAEs仅恢复9%的真实特征,表明其核心任务失败;在真实激活上,我们引入三种基线,分别将SAE特征方向或激活模式设为随机值。在多个SAE架构上的大量实验表明,这些基线在可解释性(0.87 vs 0.90)、稀疏探针(0.69 vs 0.72)和因果编辑(0.73 vs 0.72)上与完整训练的SAE表现相当。综合结果提示,当前SAE尚无法可靠分解模型内部机制。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have emerged as a promising tool for interpreting neural networks by decomposing their activations into sparse sets of human-interpretable features. Recent work has introduced multiple SAE variants and successfully scaled them to frontier models. Despite much excitement, a growing number of negative results in downstream tasks casts doubt on whether SAEs recover meaningful features. To directly investigate this, we perform two complementary evaluations. On a synthetic setup with known ground-truth features, we demonstrate that SAEs recover only $9\%$ of true features despite achieving $71\%$ explained variance, showing that they fail at their core task even when reconstruction is strong. To evaluate SAEs on real activations, we introduce three baselines that constrain SAE feature directions or their activation patterns to random values. Through extensive experiments across multiple SAE architectures, we show that our baselines match fully-trained SAEs in interpretability (0.87 vs 0.90), sparse probing (0.69 vs 0.72), and causal editing (0.73 vs 0.72). Together, these results suggest that SAEs in their current state do not reliably decompose models' internal mechanisms.

稀疏自编码器模型解释基准对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。