提出SMIXAE架构,实现语言模型中多维特征的直接发现。
SMIXAE: Towards Unsupervised Manifold Discovery in Language Models

- 用混合自编码器结构直接建模多维特征流形
- 在Gemma 2B和9B模型中成功识别已知与新流形结构
- 适合关注模型可解释性与特征发现的研究者
稀疏自编码器(SAEs)被广泛用于分解和解释神经网络激活,尤其是Transformer语言模型。然而,其主要缺陷在于无法直接建模多维特征,只能通过一组独立方向进行拼贴,需在训练后人工组合,阻碍了特征表示的可发现性与可解释性。本文提出稀疏混合自编码器(SMIXAE)架构,以解决该问题。实证表明,SMIXAE能有效直接学习先前识别出的流形结构,并在开源Gemma 2 2B与9B模型中发现新的结构。最后,讨论了当前局限并指明未来研究方向。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) have been used widely to decompose and interpret neural network activations, especially those of transformer language models. One key issue with SAEs is their inability to directly model multidimensional features. Instead, SAEs may tile such features by a set of independent directions that must be grouped together after the SAE training phase, impeding discoverability and interpretation of learned feature representations. We begin to address this issue by introducing the Sparse MIXture of Autoencoders (SMIXAE) architecture. Empirically, we provide evidence that SMIXAE models have success both in directly learning previously identified manifold structures, as well as finding novel structures, within the open source Gemma 2 2B and 9B models. Finally, we discuss several limitations and point towards areas for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。