arXiv:2410.24184cs.LG2024-10被引 2

用群交叉编码器自动发现神经网络中的对称特征,提升可解释性。

Group Crosscoders for Mechanistic Analysis of Symmetry

  • 基于对称群的字典学习,自动挖掘网络中的对称特征。
  • 在InceptionV1混合3b层中识别出具几何意义的特征家族与对称模式。
  • 适合研究模型内部表征机制的研究者使用。

我们提出群交叉编码器,作为交叉编码器的扩展,系统地发现并分析神经网络中的对称特征。尽管神经网络常在无显式结构约束下生成等变表示,但理解这些涌现对称性传统上依赖人工分析。群交叉编码器通过在对称群作用下的输入变换版本间进行字典学习,自动化该过程。应用于InceptionV1的mixed3b层,使用二面体群$\mathrm{D}_{32}$,方法揭示了若干关键见解:首先,它自然将特征聚类为可解释的家族,对应先前假设的特征类型,分离精度优于标准稀疏自编码器;其次,通过变换块分析,实现特征对称性的自动刻画,揭示曲线与直线等不同几何特征呈现不同的不变性与等变性模式。结果表明,群交叉编码器能系统揭示神经网络如何表征对称性,为机制可解释性提供新工具。

原文摘要 · Abstract (English)

We introduce group crosscoders, an extension of crosscoders that systematically discover and analyse symmetrical features in neural networks. While neural networks often develop equivariant representations without explicit architectural constraints, understanding these emergent symmetries has traditionally relied on manual analysis. Group crosscoders automate this process by performing dictionary learning across transformed versions of inputs under a symmetry group. Applied to InceptionV1's mixed3b layer using the dihedral group $\mathrm{D}_{32}$, our method reveals several key insights: First, it naturally clusters features into interpretable families that correspond to previously hypothesised feature types, providing more precise separation than standard sparse autoencoders. Second, our transform block analysis enables the automatic characterisation of feature symmetries, revealing how different geometric features (such as curves versus lines) exhibit distinct patterns of invariance and equivariance. These results demonstrate that group crosscoders can provide systematic insights into how neural networks represent symmetry, offering a promising new tool for mechanistic interpretability.

对称性可解释性神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。