稀疏自编码器的激活集无法准确反映人类对概念的典型性理解。
Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

- 用激活集合重叠度衡量相似性,比余弦相似度更可解释。
- 模型内部语义结构与人类概念边界存在明显差异。
- 适合关注模型表征机制与人类认知差异的研究者。
Shani 等人(2026)指出,大型语言模型的稠密表示虽能恢复人类类别边界,却难以体现细粒度的典型性结构。本文采用稀疏自编码器(SAE)潜在空间中活跃特征集的交集作为相似性度量,替代原有的余弦相似度。实验表明,该集合级度量在受控的模拟模型中可恢复并集式的组合结构,并在自然文本中生成语义连贯的邻域。将此方法应用于人类概念分析发现,SAE激活集未能比稠密嵌入或残差流状态更好地还原人类类别边界或类内典型性;相反,它们更忠实于模型内部的相似性结构。进一步通过可控语义扰动分析显示,人类对概念变化的判断与SAE激活集的变化之间存在显著偏差。这表明,在非理想场景下,SAE特征并非按简单的‘词袋’语义组合。
原文摘要 · Abstract (English)
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure. To probe this gap further, we study active latent sets under well-controlled semantic modifications, revealing a substantial mismatch between human judgements of conceptual change and change in the SAE active set. We interpret this as evidence that, outside idealised settings, SAE features do not compose via simple bag-of-features semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。