arXiv:2501.14926cs.LGstat.ML2025-01被引 23

通过最小化机制描述长度,分解神经网络参数以揭示内部运作原理

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition

  • 基于归因的参数分解,将模型参数拆分为忠实且简洁的机制组件
  • 在多个简化实验中成功还原了隐藏特征、压缩计算与跨层表示
  • 为理解超位置中的最小电路提供新思路,适用于各类网络结构

机械可解释性旨在理解神经网络内部学习到的机制。尽管已有进展,但如何最优地将网络参数分解为机械组件仍不明确。本文提出归因式参数分解(APD),直接将网络参数分解为三类组件:(i) 与原网络参数保持一致;(ii) 处理任意输入所需组件数量最少;(iii) 极致简单。该方法优化了网络机制的最小描述长度。我们在多个简化实验中验证了其有效性:恢复超位置中的特征、分离压缩计算过程、识别跨层分布式表示。尽管尚难扩展至非简化模型,结果为解决若干开放问题提供了方案,包括超位置中最小电路的识别、‘特征’概念的理论基础,以及一种与架构无关的神经网络分解框架。

原文摘要 · Abstract (English)

Mechanistic interpretability aims to understand the internal mechanisms learned by neural networks. Despite recent progress toward this goal, it remains unclear how best to decompose neural network parameters into mechanistic components. We introduce Attribution-based Parameter Decomposition (APD), a method that directly decomposes a neural network's parameters into components that (i) are faithful to the parameters of the original network, (ii) require a minimal number of components to process any input, and (iii) are maximally simple. Our approach thus optimizes for a minimal length description of the network's mechanisms. We demonstrate APD's effectiveness by successfully identifying ground truth mechanisms in multiple toy experimental settings: Recovering features from superposition; separating compressed computations; and identifying cross-layer distributed representations. While challenges remain to scaling APD to non-toy models, our results suggest solutions to several open problems in mechanistic interpretability, including identifying minimal circuits in superposition, offering a conceptual foundation for 'features', and providing an architecture-agnostic framework for neural network decomposition.

可解释性参数分解机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。