arXiv:2505.22506cs.LG2025-05被引 1

从表征几何角度揭示稀疏编码如何组织语言模型激活向量的结构。

Sparsification and Reconstruction from the Perspective of Representation Geometry

  • 通过分析噪声下对称半正定矩阵秩的变化,验证表征的分层结构。
  • 发现局部与全局表征能增强特征区分度,提升重建性能。
  • 为构建更可解释的稀疏自编码器提供实证依据,适合可解释性研究者。

稀疏自编码器(SAEs)已成为机制可解释性中的主流工具,旨在识别可解释的单义特征。然而,稀疏编码如何组织语言模型激活向量的表征?这种组织模式与特征解耦及重建性能有何关系?为此,本文提出SAEMA,通过观察残差流添加噪声后,沿潜在张量展开的模态张量对应的对称半正定(SSPD)矩阵的秩变化,验证了表征的分层结构。为系统研究稀疏编码对表征结构的影响,定义局部与全局表征,发现其通过合并相似语义特征并引入额外维度,增强特征间差异。进一步从优化视角干预全局表征,证明其可分性与重建性能存在显著因果关系。本研究从表征几何角度阐释稀疏性的原理,揭示表征结构变化对重建性能的影响,强调理解表征并引入表征约束的必要性,为开发新型可解释工具和改进SAEs提供实证参考。代码已公开于:https://github.com/wenjie1835/SAERepGeo。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have emerged as a predominant tool in mechanistic interpretability, aiming to identify interpretable monosemantic features. However, how does sparse encoding organize the representations of activation vector from language models? What is the relationship between this organizational paradigm and feature disentanglement as well as reconstruction performance? To address these questions, we propose the SAEMA, which validates the stratified structure of the representation by observing the variability of the rank of the symmetric semipositive definite (SSPD) matrix corresponding to the modal tensor unfolded along the latent tensor with the level of noise added to the residual stream. To systematically investigate how sparse encoding alters representational structures, we define local and global representations, demonstrating that they amplify inter-feature distinctions by merging similar semantic features and introducing additional dimensionality. Furthermore, we intervene the global representation from an optimization perspective, proving a significant causal relationship between their separability and the reconstruction performance. This study explains the principles of sparsity from the perspective of representational geometry and demonstrates the impact of changes in representational structure on reconstruction performance. Particularly emphasizes the necessity of understanding representations and incorporating representational constraints, providing empirical references for developing new interpretable tools and improving SAEs. The code is available at \hyperlink{https://github.com/wenjie1835/SAERepGeo}{https://github.com/wenjie1835/SAERepGeo}.

稀疏编码表征几何可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。