arXiv:2409.15019cs.LG2024-09被引 3

用SAE潜变量合成激活值,发现其与真实激活高度相似但仍有差异。

Evaluating Synthetic Activations composed of SAE Latents in GPT-2

  • 用SAE潜变量组合生成合成激活值,控制稀疏性和余弦相似度。
  • 合成激活与真实激活在分布上高度接近,但缺少明显的激活平台。
  • 揭示SAE潜变量具有复杂几何结构,非简单拼接可还原。

稀疏自编码器(SAEs)常用于机械可解释性研究,将残差流分解为单义的SAE潜变量。近期研究表明,早期层的激活扰动会导致模型最终层激活呈现阶跃函数式变化,且模型对真实激活与随机激活的敏感性不同。本文评估模型对合成激活的敏感性,这些合成激活由SAE潜变量构成。结果表明,在控制潜变量稀疏性和余弦相似度的前提下,合成激活与真实激活极为相似。这说明真实激活无法仅由缺乏内部结构的“SAE潜变量集合”解释,暗示SAE潜变量具有显著的几何与统计特性。值得注意的是,合成激活的激活平台效应明显弱于真实激活。

原文摘要 · Abstract (English)

Sparse Auto-Encoders (SAEs) are commonly employed in mechanistic interpretability to decompose the residual stream into monosemantic SAE latents. Recent work demonstrates that perturbing a model's activations at an early layer results in a step-function-like change in the model's final layer activations. Furthermore, the model's sensitivity to this perturbation differs between model-generated (real) activations and random activations. In our study, we assess model sensitivity in order to compare real activations to synthetic activations composed of SAE latents. Our findings indicate that synthetic activations closely resemble real activations when we control for the sparsity and cosine similarity of the constituent SAE latents. This suggests that real activations cannot be explained by a simple "bag of SAE latents" lacking internal structure, and instead suggests that SAE latents possess significant geometric and statistical properties. Notably, we observe that our synthetic activations exhibit less pronounced activation plateaus compared to those typically surrounding real activations.

可解释性SAE模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。