arXiv:2607.12166cs.LG2026-07

发现稀疏自编码器中多数可识别特征其实无实际作用,挑战传统评估方法可靠性。

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness

论文配图:From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness
图 1 · 摘自论文原文
  • 用消融与操控实验检验特征因果性,突破仅靠相似度判断的局限。
  • 高达77%的高相似特征在退化模型中无实际激活,甚至1.0相似度也无效。
  • 提出可复现的因果审计工具,适合关注模型可解释性的研究者使用。

稀疏自编码器(SAEs)是解析神经表征超叠加现象的标准工具,其评估主要依赖相关性恢复指标——即真实方向与解码器原子间的余弦相似度。我们指出该方法混淆了两个独立命题:解码器几何对齐与编码器激活行为。重现Elhage等人(2022)的超叠加相图,发现高稀疏性下存在收敛伪影,以及极端过完备状态下的未充分描述的弥散共享区。重现Gao等人(2024)的TopK与L1对比,直接验证了L1收缩效应。核心发现为因果性:对每个恢复特征进行消融与操控测试,发现在退化SAE中达77%的特征(余弦≥0.90)和良好训练模型中9%的特征具有因果惰性——即使特征存在,匹配原子也从未激活,包括余弦接近1.000的情况。我们发布saecausal-audit工具,实现模型无关、确定性流程。重新审计揭示惰性可分解为结构性惰性(反向对几何,存在于优质SAE)与竞争性惰性(退化SAE中的TopK病态);方向上则分读惰性与写惰性,五个反向对完全分离——无法监测却可通过同一原子操控,操控特异性达143-310,但消融效果为零。我们说明字节级可复现性在设计上不可行,并建议以声明堆叠形式报告,明确范围。将工具应用于生产级SAE,在小规模重现惰性模式(14%),并发现原子碰撞信号:少数原子反复成为数十个无关概念的最近匹配,跨三个批次重复出现。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) are the standard for decomposing superposed neural representations into interpretable features, and evaluation relies predominantly on correlational recovery metrics -- cosine similarity between ground-truth directions and decoder atoms. We show this conflates two distinct claims: decoder-geometry alignment and encoder-activation behavior. We reproduce the superposition phase diagram of Elhage et al. (2022), identifying a convergence artifact at high sparsity and an under-described diffuse sharing regime at extreme overcompleteness. We reproduce the TopK-versus-L1 comparison of Gao et al. (2024), with direct evidence of L1 shrinkage. Our central result is causal: subjecting every recovered feature to ablation and steering, we find up to 77% of features passing a recovery bar (cosine >= 0.90) in a degraded SAE -- and 9% in a well-trained one -- are causally inert: the matched atom never fires when the feature is present, including matches at cosine ~1.000. We package the method as sae-causal-audit, a model-agnostic instrument with a deterministic pipeline. Re-auditing refines the finding: inertness decomposes by cause into structural inertness (antipodal-pair geometry, present in good SAEs) and competitive inertness (a TopK pathology of degraded SAEs), and by direction into read- and write-inertness, which five antipodal pairs dissociate completely -- unmonitorable yet steerable through the same atom, with steering specificities of 143-310 attached to zero ablation effects. We document why byte-exact reproducibility is unavailable by construction, and propose reporting it as a stack of claims with explicit scopes. Applying the instrument to a production SAE reproduces the pattern at small scale (14% inert) and surfaces an atom-collision signal: a handful of atoms recur as the nearest match for dozens of unrelated concepts, replicated across three batches.

稀疏编码可解释性因果验证模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。