通过划分输入空间,精准定位神经网络解释的成败区域。
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction

- 按交互干预行为将输入空间分为解释良好与不足区域
- 发现高阶假设缺失的区分特征,提升解释精度
- 适合希望改进因果抽象解释的研究者使用
我们提出一种诊断神经网络解释的方法,通过识别一个高保真解释成立的输入子空间来实现。该方法特别适用于因果抽象类可解释性分析,即通过交换干预评估高层因果假设。不同于将交换干预准确率视为单一全局指标,我们通过分析成对交换干预行为,将输入空间划分为解释良好和解释不足的区域。这使因果抽象从全局评估变为诊断工具:不仅判断解释是否有效,还能揭示其在何处有效、何处失效,并识别两者差异。该诊断视角还提供实用改进策略:通过分析解释良好与不足区域的结构,可发现高阶假设中缺失的区分因素、未建模的中间变量,或将互补的部分解释整合为更强的解释。我们将其归纳为四步操作流程,在多个因果抽象场景中均产生有意义的错误分析。在小型逻辑任务中,递归应用该流程可从零开始恢复出高层假设。总体表明,输入空间划分是实现更精确、建设性与可扩展机制可解释性的关键步骤。
原文摘要 · Abstract (English)
We present a method for diagnosing interpretation in neural networks by identifying an input subspace where a proposed interpretation is highly faithful. Our method is particularly useful for causal-abstraction-style interpretability, where a high-level causal hypothesis is evaluated by interchange interventions. Rather than treating interchange intervention accuracy as a single global summary, we refine this framework by partitioning the input space into well-interpreted and under-interpreted regions according to pairwise interchange-intervention behavior. This turns causal abstraction from a purely global evaluation into a more diagnostic tool: it not only measures whether an interpretation works, but also reveals where it works, where it fails, and what distinguishes the two cases. This diagnostic view also provides practical heuristics for improving interpretations. By analyzing the structure of the well-interpreted and under-interpreted regions, we can identify missing distinctions in a high-level hypothesis, discover previously unmodeled intermediate variables, and combine complementary partial interpretations into a stronger one. We instantiate this idea as a simple four-step recipe and show that it yields informative error analyses across multiple causal abstraction settings. In a toy logic task, recursively applying the recipe recovers a high-level hypothesis from scratch. More broadly, our results suggest that partitioning the input space is a useful step toward more precise, constructive, and scalable mechanistic interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。