现有稀疏自编码器评估标准不可靠,多个核心指标失效。
Are Sparse Autoencoder Benchmarks Reliable?

- 通过重采样、合成数据和训练轨迹三角度审计评估指标
- TPP与SCR两个指标在标准设置下多次失效,不可靠
- 最可靠的sae-probes仍难区分同架构不同变体
稀疏自编码器(SAEs)是大语言模型可解释性研究的核心工具,其性能提升依赖于能有效区分优劣SAE的基准评测体系。本文对当前主流的SAEBench评测套件中的质量度量进行了三重审计:固定SAE下的重采样噪声、合成SAE上的真值相关性,以及训练轨迹间的可区分性。结果发现,两大核心指标——目标探测扰动(TPP)和虚假相关性消除(SCR)——在标准设定下均无法通过多轮检验,不应再用于SAE评估。其余指标也表现出比领域预期更高的重采样噪声和更低的可区分能力。尽管k-稀疏探测的sae-probes变体表现相对更稳定,但仍难以有效区分同一架构的不同变体。研究揭示了当前SAE评估体系存在严重缺陷,亟需建立更可靠的基准。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) are a core interpretability tool for large language models, and progress on SAE architectures depends on benchmarks that reliably distinguish better SAEs from worse ones. We audit the SAE quality metrics in SAEBench, the de-facto standard SAE evaluation suite, through three complementary lenses: reseed noise on a fixed SAE, ground-truth correlation on synthetic SAEs, and discriminability across training trajectories. We find that two of these metrics, Targeted Probe Perturbation (TPP) and Spurious Correlation Removal (SCR), fail multiple lenses at their canonical settings and should not be used to evaluate SAEs. The other metrics show higher reseed noise and lower discriminability than the field assumes. The sae-probes variant of $k$-sparse probing is the most reliable metric we tested, but even sae-probes struggles to separate variants of the same SAE architecture. Our results show the field needs better SAE benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。