arXiv:2502.17332cs.LG2025-02

提出新方法让SAE更专注模型关键特征而非简单统计。

Tokenized SAEs: Disentangling SAE Reconstructions

  • 用逐标记偏置分离词元重建与特征重建
  • 在稀疏设置下显著提升特征质量与重建效果
  • 适合研究大模型内部机制的从业者

稀疏自编码器(SAEs)已成为解析语言模型内部机制的常用工具。然而,其特征与模型中计算重要方向的对应关系尚不明确。本研究实证发现,许多RES-JB SAE特征主要对应简单的输入统计信息。我们推测这是由于训练数据存在严重类别不平衡且缺乏复杂误差信号所致。为此,提出一种将词元重建与特征重建解耦的方法:引入逐标记偏置,为有意义的重建提供增强基线。结果表明,该方法在稀疏条件下学习到更多有意义特征,并显著改善重建性能。

原文摘要 · Abstract (English)

Sparse auto-encoders (SAEs) have become a prevalent tool for interpreting language models' inner workings. However, it is unknown how tightly SAE features correspond to computationally important directions in the model. This work empirically shows that many RES-JB SAE features predominantly correspond to simple input statistics. We hypothesize this is caused by a large class imbalance in training data combined with a lack of complex error signals. To reduce this behavior, we propose a method that disentangles token reconstruction from feature reconstruction. This improvement is achieved by introducing a per-token bias, which provides an enhanced baseline for interesting reconstruction. As a result, significantly more interesting features and improved reconstruction in sparse regimes are learned.

稀疏自编码器模型解释特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。