arXiv:2603.04198stat.MLcs.LG2026-03被引 5

通过权重正则化提升稀疏自编码器的特征稳定性与可操控性。

Stable and Steerable Sparse Autoencoders with Weight Regularization

  • 引入L2权重正则化,结合权值共享与单位范数约束,增强特征一致性。
  • 在语言模型上,跨种子共享特征比例提升,控制成功率翻倍。
  • 正则化后解释性评分更好预测功能可控性,利于人工理解。

稀疏自编码器(SAEs)广泛用于提取神经网络激活中的可解释特征,但其学习到的特征易受随机种子和训练设置影响。为提升稳定性,我们研究了对编码器和解码器权重添加L1或L2正则化的效果,并评估其与常见训练默认设置的交互。在MNIST上,L2权重正则化产生一组高度一致的核心特征;当结合权值共享初始化与单位范数解码器约束时,跨种子特征一致性显著提升。在语言模型(Pythia-70M-deduped)的TopK SAE训练中,加入小量L2权重惩罚使三个随机种子间共享特征比例增加,控制成功率大致翻倍,而自动可解释性分数均值基本不变。此外,在正则化条件下,激活控制成功率更可由自动可解释性评分预测,表明正则化使文本解释与功能可控性对齐。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) are widely used to extract human-interpretable features from neural network activations, but their learned features can vary substantially across random seeds and training choices. To improve stability, we studied weight regularization by adding L1 or L2 penalties on encoder and decoder weights, and evaluate how regularization interacts with common SAE training defaults. On MNIST, we observe that L2 weight regularization produces a core of highly aligned features and, when combined with tied initialization and unit-norm decoder constraints, it dramatically increases cross-seed feature consistency. For TopK SAEs trained on language model activations (Pythia-70M-deduped), adding a small L2 weight penalty increased the fraction of features shared across three random seeds and roughly doubles steering success rates, while leaving the mean of automated interpretability scores essentially unchanged. Finally, in the regularized setting, activation steering success becomes better predicted by auto-interpretability scores, suggesting that regularization can align text-based feature explanations with functional controllability.

稀疏自编码器特征可解释性权重正则化语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。