arXiv:2511.04638cs.LGcs.AI2025-11被引 5

干预神经网络表征可能偏离原始分布,影响解释可靠性。

Addressing divergent representations from causal interventions on neural networks

  • 通过因果干预分析模型内部表征的分布变化
  • 发现干预常导致表征偏离自然分布,分无害与有害两类
  • 改进反事实潜在损失以减少有害偏差,提升解释可信度

机制可解释性常用因果干预手段操纵模型表征以理解其编码内容。本文探讨此类干预是否会产生分布外(发散)的表征,并质疑其解释对原模型状态的忠实性。理论与实证表明,常见干预方法常使内部表征偏离目标模型的自然分布。进一步分析两类发散:在层行为零空间中的‘无害’发散,以及激活隐藏路径并引发潜在行为改变的‘有害’发散。为缓解后者,我们应用并改进Grant(2025)提出的反事实潜在(CL)损失,使干预后表征更贴近自然分布,降低有害发散风险,同时保持干预的可解释能力。结果为更可靠的可解释方法指明方向。

原文摘要 · Abstract (English)

A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to understand what those representations encode. Here we ask whether such interventions create out-of-distribution (divergent) representations, and whether this raises concerns about how faithful their resulting explanations are to the target model in its natural state. First, we demonstrate theoretically and empirically that common causal intervention techniques often do shift internal representations away from the natural distribution of the target model. Then, we provide a theoretical analysis of two cases of such divergences: "harmless" divergences that occur in the behavioral null-space of the layer(s) of interest, and "pernicious" divergences that activate hidden network pathways and cause dormant behavioral changes. Finally, in an effort to mitigate the pernicious cases, we apply and modify the Counterfactual Latent (CL) loss from Grant (2025) allowing representations from causal interventions to remain closer to the natural distribution, reducing the likelihood of harmful divergences while preserving the interpretive power of the interventions. Together, these results highlight a path towards more reliable interpretability methods.

可解释性因果干预表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。