arXiv:2606.27510cs.LGcs.CL2026-06被引 6

激活修补法隐藏了组件间的交互效应,可能误导对模型机制的理解。

The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

  • 通过因果中介分析重新推导修补法,发现其结果包含组件间交互项。
  • 交互效应随激活差异增大而增强,且解释了忠诚度评分的不稳定性。
  • 交互项是诊断模型依赖性的重要信号,提示需组合搜索才能发现机制。

激活修补法是机制可解释性中的主要工具,通过估计自然间接效应(NIE)将模型行为的因果责任归因于各个组件。从因果中介分析重新推导后发现,NIE不仅反映特定组件的因果效应,还包含交互效应(INT),即该组件的效应如何依赖于其他组件的状态。尝试消除INT的修正方法均有预设失败模式:在GPT-2 IOI电路中,依赖其他组件状态的组件或被忽略,或被人为夸大;交互效应方差解释了此前记录的忠诚度分数不稳定性。我们证明,INT与干净与修补后组件激活值之间的距离成正比,在局部仿射模型中可忽略,并可分解为成对及更高阶的组合交互。尽管无法避免,但INT并非需要消除的干扰,而是可解释性研究的诊断工具——其个体和群体层面的大小与符号,能揭示因果结论是否依赖于提示,以及基于贪婪NIE排序会遗漏哪些仅通过组合搜索才能发现的机制。

原文摘要 · Abstract (English)

Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains interaction effects (INT) that measure how much the component's causal effect itself depends on the state of other components in the model. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes. We demonstrate these failure modes in the GPT-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher-order group interactions. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies. Its individual and group-level magnitude and sign signal when causal conclusions are prompt-dependent, and when greedy NIE-based component ranking will miss mechanisms only discoverable through combinatorial search.

可解释性因果推理模型机制交互效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。