arXiv:2603.09988cs.CLcs.AI2026-03

让大模型的内部机制能用自然语言说清楚,且解释可信。

Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations

  • 用激活修补法找出关键注意力头,连接电路分析与语言解释。
  • 解释准确率达100%充分性,但仅22%全面性,发现备用机制存在。
  • 大模型自动生成解释比模板好64%,适合想理解模型决策的人。

机制可解释性旨在识别模型行为背后的内部电路,但将这些发现转化为人类可理解的解释仍是难题。我们提出一个流程:(i) 通过激活修补法识别因果上重要的注意力头;(ii) 使用模板和大模型生成解释;(iii) 采用适配电路级归因的ERASER风格指标评估解释的忠实度。在GPT-2 Small(124M参数)的间接宾语识别(IOI)任务上,我们识别出6个注意力头,贡献了61.4%的logit差异。基于电路的解释达到100%充分性,但仅有22%全面性,揭示出分布式备份机制。大模型生成的解释在质量指标上比模板基线高64%。我们发现模型置信度与解释忠实度之间无相关性(r = 0.009),并识别出三类导致解释偏离机制的失败模式。

原文摘要 · Abstract (English)

Mechanistic interpretability identifies internal circuits responsible for model behaviors, yet translating these findings into human-understandable explanations remains an open problem. We present a pipeline that bridges circuit-level analysis and natural language explanations by (i) identifying causally important attention heads via activation patching, (ii) generating explanations using both template-based and LLM-based methods, and (iii) evaluating faithfulness using ERASER-style metrics adapted for circuit-level attribution. We evaluate on the Indirect Object Identification (IOI) task in GPT-2 Small (124M parameters), identifying six attention heads accounting for 61.4% of the logit difference. Our circuit-based explanations achieve 100% sufficiency but only 22% comprehensiveness, revealing distributed backup mechanisms. LLM-generated explanations outperform template baselines by 64% on quality metrics. We find no correlation (r = 0.009) between model confidence and explanation faithfulness, and identify three failure categories explaining when explanations diverge from mechanisms.

可解释性大模型注意力头自然语言解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。