通过随机遮蔽提升稀疏自编码器的鲁棒性,减少特征吸收问题。
Improving Robustness In Sparse Autoencoders via Masked Regularization

- 训练时随机遮蔽输入令牌,打破共现模式。
- 降低特征吸收率,提升探针性能与分布外泛化能力。
- 适用于各类稀疏自编码器架构,增强可解释性工具可靠性。
稀疏自编码器(SAEs)广泛用于大语言模型激活的机制可解释性分析,将激活投影到稀疏潜在空间。然而,仅靠稀疏性无法充分代表可解释性,现有训练目标常导致潜在表示脆弱。SAEs易发生特征吸收现象——由于共现,通用特征被更具体特征覆盖,导致可解释性下降,尽管重建保真度高。近期关于分布外(OOD)性能的负面结果进一步揭示了训练目标不明确引发的鲁棒性缺陷。为此,我们提出一种基于遮蔽的正则化方法:在训练中随机替换令牌,以破坏共现模式。该方法提升了多种SAE架构和稀疏度下的鲁棒性,减少了特征吸收,增强了探针性能,并缩小了分布外差距。结果表明,这为构建更可靠的可解释性工具提供了可行路径。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) are widely used in mechanistic interpretability to project LLM activations onto sparse latent spaces. However, sparsity alone is an imperfect proxy for interpretability, and current training objectives often result in brittle latent representations. SAEs are known to be prone to feature absorption, where general features are subsumed by more specific ones due to co-occurrence, degrading interpretability despite high reconstruction fidelity. Recent negative results on Out-of-Distribution (OOD) performance further underscore broader robustness related failures tied to under-specified training objectives. We address this by proposing a masking-based regularization that randomly replaces tokens during training to disrupt co-occurrence patterns. This improves robustness across SAE architectures and sparsity levels reducing absorption, enhancing probing performance, and narrowing the OOD gap. Our results point toward a practical path for more reliable interpretability tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。