arXiv:2504.00194cs.LGcs.AI2025-04被引 5

通过参数空间分解,精准定位模型中稀疏激活的计算电路。

Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition

  • 基于损失曲面局部分解,寻找可重构梯度的方向子集。
  • 在模拟模型中几乎完美恢复预设的低秩子网络。
  • 适用于真实Transformer与CNN,适合关注模型内部机制的研究者。

现有神经网络可解释性研究多聚焦于激活空间,但难以揭示底层计算电路。为此,我们提出一种新方法——局部损失曲面分解(L3D),旨在识别参数空间中的低秩子网络:即某些参数方向的线性组合,能重建任意样本输出与参考输出间的损失梯度。我们设计了一系列具有明确子网络结构的简化模型,验证了L3D能近乎完美地恢复这些子网络。进一步分析发现,沿特定子网络方向扰动模型,仅影响相关样本。最后,我们将L3D应用于真实Transformer与卷积神经网络,证明其在参数空间中识别可解释且相关的计算电路的潜力。

原文摘要 · Abstract (English)

Much of mechanistic interpretability has focused on understanding the activation spaces of large neural networks. However, activation space-based approaches reveal little about the underlying circuitry used to compute features. To better understand the circuits employed by models, we introduce a new decomposition method called Local Loss Landscape Decomposition (L3D). L3D identifies a set of low-rank subnetworks: directions in parameter space of which a subset can reconstruct the gradient of the loss between any sample's output and a reference output vector. We design a series of progressively more challenging toy models with well-defined subnetworks and show that L3D can nearly perfectly recover the associated subnetworks. Additionally, we investigate the extent to which perturbing the model in the direction of a given subnetwork affects only the relevant subset of samples. Finally, we apply L3D to a real-world transformer model and a convolutional neural network, demonstrating its potential to identify interpretable and relevant circuits in parameter space.

可解释性神经网络参数空间电路识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。