arXiv:2607.20596cs.LGcs.CL2026-07被引 1

检验单标记稀疏自编码器特征的因果必要性,发现其作用因模型层深和训练方法而异。

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

论文配图:Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
图 1 · 摘自论文原文
  • 通过零消融测试390万特征,验证单标记特征在解码空间中聚集于早期层且更紧密。
  • 删除一个特征使208种层条件下178种出现目标词逻辑值下降,显著影响模型输出。
  • 不同训练方法导致特征因果性差异,跨家族效果大于同一模型内规模变化。

稀疏自编码器(SAE)特征被用于解释和引导大语言模型,但尚未有人测试其因果角色在不同SAE族间是否稳定。单标记特征仅对一个词汇项激活,便于直接比较。我们分析了六种模型与三种SAE族共390万特征,并在完整层深度下进行零消融:这些特征在解码空间中聚集得更紧密(4.7倍紧致),主要集中在早期层。删除任一特征会使模型在208种层配置中的178种情况下该词的逻辑值下降,经多重比较校正后仍显著。层深决定损伤传递方式:早期层删除影响后续层,晚期层删除直接影响输出。跨族因果差异超过族内规模效应:同一基础模型上,GemmaScope与BatchTopK特征具因果锚定性,而LlamaScope特征局部冗余;在LlamaScope中,该词排名在消融后恢复至原始2倍内的概率为96%-98%。仅改变激活函数即可反转此差异符号,说明训练方案是剩余关键因素:跨族因果结论对训练方法敏感,不仅取决于激活函数或规模。

原文摘要 · Abstract (English)

Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet nobody has tested whether a feature's causal role is stable across SAE families. Single-token features fire on one vocabulary item, so ground truth permits direct comparison. We analyze 3.9M features across six models and three SAE families and zero-ablate at full layer depth: they sit 4.7x tighter in decoder space and concentrate in early layers. Deleting one lowers the model's logit for that token in 178 of 208 layer conditions, significant after multiple-comparison correction. And depth decides how the damage lands: early-layer deletions disrupt the layers that follow, late-layer deletions change the output directly. Cross-family causal differences exceed within-family scale effects: on the same base model, GemmaScope and BatchTopK features are causally anchored, LlamaScope features locally redundant. Under LlamaScope the token returns to within 2x its pre-ablation rank 96-98% of the time. Changing only the activation function reverses the sign of that difference, so the training recipe is the remaining candidate: cross-family claims are sensitive to training methodology, not just activation function or scale.

稀疏自编码器因果分析大模型解释层深度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。