arXiv:2607.08349cs.LG2026-07中稿 · UAI 2026被引 1

为模型解释的干预实验提供可认证的统计保障,避免误判结果真假。

Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability

  • 将解释效果建模为因果效应,用置信区间和实时验证序列量化可靠性。
  • 在MNIST和GPT-2上验证高保真性,识别出不显著的方法差异。
  • 支持自适应采样,降低10~30倍计算成本,适合严谨评估解释可信度的研究者。

机制可解释性常通过干预模型(如替换隐藏状态、激活值修补、组件消融或压缩模型对比)来评估解释效果,但这类实验通常只报告点估计,即使评估过程可监控或自适应调整。这使得难以判断报告的保真度或修补效果是稳定因果结论,还是有限采样与评估选择的结果。本文提出认证干预保真度(CIF),一种用于干预式可解释性评估的统计层。CIF首先将报告量定义为因果估计量:在指定输入分布与干预分布上的有界评分期望。随后,它提供置信区间与任意时间有效的置信序列,包括通过有界混合重要性加权实现的自适应干预采样。我们实例化了基于霍夫丁的序列与方差自适应投注序列,后者在实验中将认证成本降低了10~30倍。在MNIST抽象任务与GPT-2 Small IOI电路上,CIF能认证高保真性结论,揭示表面方法差异并无统计支持,并显式展现对干预分布的敏感性。

原文摘要 · Abstract (English)

Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one. These experiments are usually summarized by a point estimate, even though the evaluation may be monitored while it runs or adapted toward suspected failures. This makes it hard to tell whether a reported fidelity or patching effect is a stable causal claim or a consequence of finite sampling and evaluation choices. We introduce Certified Interventional Fidelity (CIF), a statistical layer for interventional interpretability evaluations. CIF first writes the quantity being reported as a causal estimand: an expectation of a bounded score over a stated input distribution and a stated intervention distribution. It then provides confidence intervals and anytime-valid confidence sequences for this estimand, including under adaptive intervention sampling via bounded mixture importance weighting. We instantiate CIF with Hoeffding-style sequences and variance-adaptive betting sequences, the latter reducing certification cost by 10-30x in our experiments. On MNIST abstractions and GPT-2 Small IOI circuits, CIF certifies high-fidelity claims, shows when apparent method differences are not statistically supported, and makes sensitivity to the intervention distribution explicit.

可解释性因果推断统计验证模型干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。