通过参数空间分解,精准定位模型中稀疏激活的计算电路。
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
- 基于损失曲面局部分解,寻找可重构梯度的方向子集。
- 在模拟模型中几乎完美恢复预设的低秩子网络。
- 适用于真实Transformer与CNN,适合关注模型内部机制的研究者。
现有神经网络可解释性研究多聚焦于激活空间,但难以揭示底层计算电路。为此,我们提出一种新方法——局部损失曲面分解(L3D),旨在识别参数空间中的低秩子网络:即某些参数方向的线性组合,能重建任意样本输出与参考输出间的损失梯度。我们设计了一系列具有明确子网络结构的简化模型,验证了L3D能近乎完美地恢复这些子网络。进一步分析发现,沿特定子网络方向扰动模型,仅影响相关样本。最后,我们将L3D应用于真实Transformer与卷积神经网络,证明其在参数空间中识别可解释且相关的计算电路的潜力。
原文摘要 · Abstract (English)
Much of mechanistic interpretability has focused on understanding the activation spaces of large neural networks. However, activation space-based approaches reveal little about the underlying circuitry used to compute features. To better understand the circuits employed by models, we introduce a new decomposition method called Local Loss Landscape Decomposition (L3D). L3D identifies a set of low-rank subnetworks: directions in parameter space of which a subset can reconstruct the gradient of the loss between any sample's output and a reference output vector. We design a series of progressively more challenging toy models with well-defined subnetworks and show that L3D can nearly perfectly recover the associated subnetworks. Additionally, we investigate the extent to which perturbing the model in the direction of a given subnetwork affects only the relevant subset of samples. Finally, we apply L3D to a real-world transformer model and a convolutional neural network, demonstrating its potential to identify interpretable and relevant circuits in parameter space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。