arXiv:2508.15094cs.LG2025-08Conference of the …被引 4

用新指标证明稀疏自编码器能显著提升语言模型的语义可解释性。

Evaluating Sparse Autoencoders for Monosemantic Representation

  • 提出基于JS散度的细粒度概念分离度评分,量化神经元激活分布差异。
  • 在两个大模型上验证:SAEs使概念分离度提升23%~41%,减少多义性。
  • 提出新干预方法APP,部分抑制时仅增1.8%困惑度却有效移除概念。

大语言模型的解释性受限于神经元的多义性,即一个神经元对多个无关概念激活。稀疏自编码器(SAEs)通过将密集激活转换为稀疏、更易解读的特征,被提出以缓解此问题。尽管已有研究认为SAEs促进单义性,但尚无定量比较分析其与基线模型在概念激活分布上的差异。本文首次通过激活分布视角系统评估SAEs与基线模型。我们引入基于杰恩-锡尔顿距离(Jensen-Shannon distance)的细粒度概念分离度评分,衡量神经元在不同概念下的激活分布差异。在Gemma-2-2B和DeepSeek-R1两个大模型上,使用五个数据集(包括词级与句级),我们发现SAEs显著降低多义性,概念分离度提升23%~41%。为评估实用性,我们测试了两种概念干预策略:全神经元掩码与部分抑制。结果表明,在部分抑制下,相较于基线模型,SAEs实现更精准的概念控制。基于此,我们提出基于后验概率的衰减方法(Attenuation via Posterior Probabilities, APP),利用条件激活分布进行目标抑制。APP在保持高效概念移除的同时,仅导致1.8%的困惑度增加,优于现有方法。

原文摘要 · Abstract (English)

A key barrier to interpreting large language models is polysemanticity, where neurons activate for multiple unrelated concepts. Sparse autoencoders (SAEs) have been proposed to mitigate this issue by transforming dense activations into sparse, more interpretable features. While prior work suggests that SAEs promote monosemanticity, no quantitative comparison has examined how concept activation distributions differ between SAEs and their base models. This paper provides the first systematic evaluation of SAEs against base models through activation distribution lens. We introduce a fine-grained concept separability score based on the Jensen-Shannon distance, which captures how distinctly a neuron's activation distributions vary across concepts. Using two large language models (Gemma-2-2B and DeepSeek-R1) and multiple SAE variants across five datasets (including word-level and sentence-level), we show that SAEs reduce polysemanticity and achieve higher concept separability. To assess practical utility, we evaluate concept-level interventions using two strategies: full neuron masking and partial suppression. We find that, compared to base models, SAEs enable more precise concept-level control when using partial suppression. Building on this, we propose Attenuation via Posterior Probabilities (APP), a new intervention method that uses concept-conditioned activation distributions for targeted suppression. APP achieves the smallest perplexity increase while remaining highly effective at concept removal.

稀疏编码模型解释概念控制自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。