改进敏感方向实验基线,揭示SAE特征对模型输出的影响规律。
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
- 提出新基线提升扰动方向评估可靠性
- 低稀疏度SAE特征影响更大,效果显著
- 端到端SAE不优于传统SAE,适合对比研究
敏感方向实验通过扰动激活值沿特定方向,测量其对下一个词预测概率的影响,以理解语言模型的计算特性。本文引入一种改进的扰动方向基线,发现稀疏自编码器(SAE)重构误差的KL散度不再异常偏高。结果表明,不同稀疏度下的SAE特征方向对模型输出的影响存在差异,稀疏度更低(即L0更小)的特征方向影响更强。此外,端到端训练的SAE特征并未表现出比传统SAE更强的输出影响,为特征可解释性研究提供了更可靠的比较基准。
原文摘要 · Abstract (English)
Sensitive directions experiments attempt to understand the computational features of Language Models (LMs) by measuring how much the next token prediction probabilities change by perturbing activations along specific directions. We extend the sensitive directions work by introducing an improved baseline for perturbation directions. We demonstrate that KL divergence for Sparse Autoencoder (SAE) reconstruction errors are no longer pathologically high compared to the improved baseline. We also show that feature directions uncovered by SAEs have varying impacts on model outputs depending on the SAE's sparsity, with lower L0 SAE feature directions exerting a greater influence. Additionally, we find that end-to-end SAE features do not exhibit stronger effects on model outputs compared to traditional SAEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。