arXiv:2506.15963cs.LG2025-06中稿 · ICLR被引 14

揭示稀疏自编码器恢复语义特征的理论极限并提出改进方法

On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy

  • 首次给出稀疏自编码器的闭式解,发现仅当语义特征极稀疏时才能完全恢复
  • 提出加权策略,显著提升一般情况下特征的单义性和可解释性
  • 理论指导权重选择,适用于需要理解大模型内部表征的研究者

稀疏自编码器(SAEs)近年来成为解析大语言模型(LLMs)学习特征的强大工具。通过用稀疏激活网络重建特征,SAEs旨在将复杂的叠加多义特征分解为可解释的单义特征。尽管应用广泛,但尚不清楚在何种条件下SAEs能从叠加的多义特征中完全恢复出真实单义特征。本文首次提供SAEs的闭式解理论分析,揭示其通常无法完全恢复真实单义特征,除非真实特征极度稀疏。为改善一般情况下的特征恢复效果,我们提出一种重加权策略,聚焦于增强真实单义特征的重建而非观测到的多义特征。我们进一步建立了所提加权自编码器(WSAE)的理论权重选择原则。多组实验验证了理论结论,并表明我们的WSAE显著提升了特征的单义性和可解释性。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have recently emerged as a powerful tool for interpreting the features learned by large language models (LLMs). By reconstructing features with sparsely activated networks, SAEs aim to recover complex superposed polysemantic features into interpretable monosemantic ones. Despite their wide applications, it remains unclear under what conditions SAEs can fully recover the ground truth monosemantic features from the superposed polysemantic ones. In this paper, we provide the first theoretical analysis with a closed-form solution for SAEs, revealing that they generally fail to fully recover the ground truth monosemantic features unless the ground truth features are extremely sparse. To improve the feature recovery of SAEs in general cases, we propose a reweighting strategy targeting at enhancing the reconstruction of the ground truth monosemantic features instead of the observed polysemantic ones. We further establish a theoretical weight selection principle for our proposed weighted SAE (WSAE). Experiments across multiple settings validate our theoretical findings and demonstrate that our WSAE significantly improves feature monosemanticity and interpretability.

稀疏自编码器特征可解释性大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。