用生物启发的算法在CLIP表示中发现分层单义神经元。
Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm

- 基于类内对比和类别路由的前向-前向训练,不依赖稀疏性约束。
- 在CLIP特征上训练出抽象层级递增的单义神经元,捕捉独立于前景的环境属性。
- 无需监督抽象层次,可从零训练网络,性能领先同类前向算法。
机制可解释性在理解神经网络表征方面取得显著进展,其中稀疏字典学习(SDL)方法,尤其是稀疏自编码器,是核心范式。然而,近期研究指出该范式的若干局限:SDL目标不可识别;依赖线性表征假设;越来越多证据表明概念以非线性方式编码,无法表示为单一方向。我们提出另一种通向单义性的路径。生物视觉系统中的神经元高度选择性,并形成抽象层级递进的结构,这种组织由局部、逐层的学习规则实现,而非全局误差信号。因此我们探究:是否一种生物合理的学习算法也能产生单义神经元?为此,我们提出组对比前向-前向(GCFF)算法,结合类别特异性路由与类内对比目标,通过架构约束实现单义性,而非稀疏性。由于GCFF在待研究表征上附加多个非线性层,其神经元可捕捉非线性概念。在CLIP表征上,单个训练后的GCFF模块恢复了抽象程度随深度递增的单义神经元,能识别独立于图像前景的环境属性,且无需稀疏性约束或抽象层次监督。我们进一步证明,GCFF可从零训练网络,在多个图像分类基准上达到前向-前向算法中的最先进水平。
原文摘要 · Abstract (English)
Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse autoencoders, as a central paradigm. However, recent work has reported several limitations of this paradigm: SDL objectives are non-identifiable; SDL methods rely heavily on the Linear Representation Hypothesis; and a growing body of evidence points to concepts that are encoded non-linearly and are therefore not expressible as any single direction. We hypothesise that a different route to monosemanticity is available. Biological visual systems exhibit highly selective neurons organised into hierarchies of increasing abstraction, and this organisation emerges from local, layer-wise learning rules rather than from a global error signal; we therefore ask whether a biologically plausible learning algorithm will likewise yield monosemantic neurons. To test this, we propose Group-Contrastive Forward-Forward (GCFF), a forward-forward training algorithm that combines class-specific routing with within-class contrastive objectives, reaching monosemanticity through architectural constraints rather than sparsity. Because GCFF attaches multiple non-linear layers to the representation under study, its neurons can therefore capture the non-linear concepts. On CLIP representations, a single trained GCFF module recovers monosemantic neurons whose abstraction increases progressively with depth, reaching environmental properties that hold independently of an image's foreground, without any sparsity constraint or supervision of abstraction level. We further demonstrate that GCFF can train networks from scratch, achieving state-of-the-art performance among forward-forward algorithms on various image classification benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。