提出可衡量神经网络可解释性方法真实进步的基准测试
MIB: A Mechanistic Interpretability Benchmark
- 构建双轨评估框架,分别检验模型组件定位和因果变量识别能力
- 发现归因与掩码优化法在电路定位上表现最优,监督式DAS在变量定位中领先
- 揭示稀疏自编码器特征不优于原始神经元,为方法选择提供实证依据
如何判断新的机制可解释性方法是否真正取得进展?为此,我们提出MIB(机制可解释性基准),包含两个赛道、四项任务和五种模型。MIB侧重于方法能否精确且简洁地恢复神经语言模型中的相关因果路径或因果变量。电路定位赛道比较不同方法对执行任务至关重要的模型组件及其连接的定位能力(如归因修补或信息流路径);因果变量定位赛道比较将隐藏向量特征化(如稀疏自编码器SAE或分布式对齐搜索DAS)并将其与任务相关的因果变量对齐的能力。使用MIB发现,归因与掩码优化方法在电路定位中表现最佳;在因果变量定位中,监督式DAS方法最优,而SAE特征并不优于原始神经元(即未特征化的隐藏向量)。这些结果表明MIB能够实现有意义的比较,增强我们对领域内确实存在实质性进展的信心。
原文摘要 · Abstract (English)
How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recover relevant causal pathways or causal variables in neural language models. The circuit localization track compares methods that locate the model components - and connections between them - most important for performing a task (e.g., attribution patching or information flow routes). The causal variable localization track compares methods that featurize a hidden vector, e.g., sparse autoencoders (SAEs) or distributed alignment search (DAS), and align those features to a task-relevant causal variable. Using MIB, we find that attribution and mask optimization methods perform best on circuit localization. For causal variable localization, we find that the supervised DAS method performs best, while SAE features are not better than neurons, i.e., non-featurized hidden vectors. These findings illustrate that MIB enables meaningful comparisons, and increases our confidence that there has been real progress in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。