用激活差异聚类提取可解释概念,效果接近有监督方法。
The Deleuzian Representation Hypothesis
- 通过聚类激活差异提取概念,基于判别分析理论
- 在五模型三模态上表现优于无监督SAE,接近有监督基线
- 提取的概念能操控模型内部表示,适合可解释性研究
我们提出一种替代稀疏自编码器(SAEs)的简单有效无监督方法,用于从神经网络中提取可解释概念。核心思想是聚类激活差异,并在判别分析框架下给出形式化证明。为提升概念多样性,通过激活偏度加权优化聚类过程。该方法契合德勒兹关于概念即差异的现代观点。我们在五个模型和三种模态(视觉、语言、音频)上评估,衡量概念质量、多样性和一致性。结果表明,该方法在概念质量上超越现有无监督SAE变体,接近有监督基线,且提取的概念可实现对模型内部表示的控制,证明其对下游行为具有因果影响。
原文摘要 · Abstract (English)
We propose an alternative to sparse autoencoders (SAEs) as a simple and effective unsupervised method for extracting interpretable concepts from neural networks. The core idea is to cluster differences in activations, which we formally justify within a discriminant analysis framework. To enhance the diversity of extracted concepts, we refine the approach by weighting the clustering using the skewness of activations. The method aligns with Deleuze's modern view of concepts as differences. We evaluate the approach across five models and three modalities (vision, language, and audio), measuring concept quality, diversity, and consistency. Our results show that the proposed method achieves concept quality surpassing prior unsupervised SAE variants while approaching supervised baselines, and that the extracted concepts enable steering of a model's inner representations, demonstrating their causal influence on downstream behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。