arXiv:2502.02470cs.LGcs.AI2025-02被引 2

通过模块化损失训练神经网络,使其形成独立小电路。

Studying Cross-cluster Modularity in Neural Networks

  • 引入模块化损失函数,让神经网络自动形成互不干扰的子模块。
  • 训练后的模型电路更小,但任务专属性未提升。
  • 适用于想理解模型内部结构的研究者或可解释性开发者。

提升神经网络可解释性的方法之一是增强其可聚类性,即把模型拆分为互不重叠的独立模块进行研究。本文定义了一种可聚类性度量,并发现预训练模型通过谱图聚类形成高度纠缠的簇。为此,我们引入一种“模块化损失”来训练模型,促使其形成非交互的簇。随后研究了这类高度聚类模型的特性:结果表明,这些模型虽未表现出更强的任务特异性,但形成了更小的内部电路。实验涵盖在MNIST和CIFAR上训练的CNN、在模加法任务上的小型Transformer,以及在Wiki数据集上的GPT-2和Pythia,还有在化学数据集上的Gemma。该研究揭示了聚类模型可能具备的典型特征。

原文摘要 · Abstract (English)

An approach to improve neural network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We define a measure for clusterability and show that pre-trained models form highly enmeshed clusters via spectral graph clustering. We thus train models to be more modular using a "clusterability loss" function that encourages the formation of non-interacting clusters. We then investigate the emerging properties of these highly clustered models. We find our trained clustered models do not exhibit more task specialization, but do form smaller circuits. We investigate CNNs trained on MNIST and CIFAR, small transformers trained on modular addition, and GPT-2 and Pythia on the Wiki dataset, and Gemma on a Chemistry dataset. This investigation shows what to expect from clustered models.

可解释性模块化神经网络结构聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。