arXiv:2506.15679cs.LGcs.AI2025-06NeurIPS被引 17

dense latent 是语言模型的有用特征,非训练噪声。

Dense SAE Latents Are Features, Not Bugs

  • 发现 dense latents 呈反向成对分布,重构残差流特定方向。
  • 识别出位置追踪、词性标记等 6 类功能型 dense latent,具明确语义。
  • 从早期到末层,dense features 由结构特征演变为输出信号,具层次性。

稀疏自编码器(SAEs)旨在通过稀疏约束提取语言模型中的可解释特征。然而,许多 SAE latent 激活频繁(即密集),引发其是否为训练伪影的担忧。本文系统研究了 dense latents 的几何结构、功能与起源,发现它们不仅持续存在,且常反映有意义的模型表征。我们首先证明,dense latents 常成反向对出现,重构残差流中特定方向;删除其子空间后,重训练 SAE 时新 dense features 不再涌现,表明高密度特征是残差空间的内在属性。接着提出 dense latents 分类体系,涵盖位置追踪、上下文绑定、熵调控、字母特异性输出信号、词性标注及主成分重建等类别。最后分析其跨层演化:早期层以结构特征为主,中层转向语义特征,末层则聚焦输出导向信号。结果表明,dense latents 在语言模型计算中具有功能性作用,不应视为训练噪声。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) are designed to extract interpretable features from language models by enforcing a sparsity constraint. Ideally, training an SAE would yield latents that are both sparse and semantically meaningful. However, many SAE latents activate frequently (i.e., are \emph{dense}), raising concerns that they may be undesirable artifacts of the training procedure. In this work, we systematically investigate the geometry, function, and origin of dense latents and show that they are not only persistent but often reflect meaningful model representations. We first demonstrate that dense latents tend to form antipodal pairs that reconstruct specific directions in the residual stream, and that ablating their subspace suppresses the emergence of new dense features in retrained SAEs -- suggesting that high density features are an intrinsic property of the residual space. We then introduce a taxonomy of dense latents, identifying classes tied to position tracking, context binding, entropy regulation, letter-specific output signals, part-of-speech, and principal component reconstruction. Finally, we analyze how these features evolve across layers, revealing a shift from structural features in early layers, to semantic features in mid layers, and finally to output-oriented signals in the last layers of the model. Our findings indicate that dense latents serve functional roles in language model computation and should not be dismissed as training noise.

自编码器语言模型特征表示稠密激活

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。