从几何视角解析稀疏自编码器中概念学习与神经元解释的机制
A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders

- 将概念定义为数据点集合,建模为人类与模型概念的对齐问题
- 提出检测、分离、近似三层次学习,给出可表示概念的几何条件和误差界
- 揭示特征分裂、吸收等现象的集合论本质,适合关注可解释性研究者
我们提出了一个统一的数学框架,用于从几何角度理解稀疏自编码器(SAEs)中的概念学习与神经元解释。尽管SAE通过学习稀疏特征表示提升了神经网络的可解释性,但“概念”和“学习”的严格定义仍不清晰。我们将概念形式化为数据点集合,并将概念学习建模为人类定义与模型生成概念之间的集合对齐问题。该框架区分了三种逐步增强的学习层次——检测、分离与近似,并推导出概念由单个神经元或多神经元单元表示时的几何条件、误差边界和容量约束。同时,该理论为常见的SAE现象(如特征分裂、特征吸收、特征族、层次概念)提供了集合论解释。最后,通过形式概念分析连接概念学习与神经元解释,指出二者不必一致,且其多对多关系可通过概念格组织。在使用ReLU和Top-K SAE的合成数据上的实验验证了理论,并揭示了SAE规模与稀疏度对概念学习的影响。
原文摘要 · Abstract (English)
We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpretability of neural networks by learning sparse feature representations, a principled definition of ''concept'' and ''learning'' remains unclear. We formalize concepts as sets of data points and cast concept learning as a set-alignment problem between human-defined and model-induced concepts. This formulation distinguishes three increasingly strong notions of learning -- detection, separation, and approximation -- and yields geometric conditions, error bounds, and capacity constraints for when concepts can be represented by individual neurons or multi-neuron units. It also provides a set-theoretic account for common SAE phenomena, including feature splitting, feature absorption, feature families, and hierarchical concepts. Finally, we connect concept learning and neuron interpretation through formal concept analysis, showing that the two directions need not agree and that their many-to-many structure can be organized by concept lattices. Experiments on synthetic data with ReLU and Top-$K$ SAEs illustrate the theory and reveal the effects of SAE size and sparsity on concept learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。