arXiv:2605.08012cs.LGcs.AI2026-05

解释神经网络时需明示因果推断的假设,否则结论不可靠

Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims

  • 指出机械可解释性研究常误用因果术语,却未说明关键假设
  • 审计30篇论文发现90%未设专门识别假设章节,验证指标被当作因果证据
  • 建议建立披露规范:明确因果声明、列出假设并说明其失效影响

机械可解释性研究日益使用因果词汇(如电路、中介、因果抽象、单义性),但这些主张需依赖明确的识别假设。对四类方法共10篇论文的有目的性审计显示,均无专门的识别假设章节,且验证指标(如忠实性、完整性、单义性、对齐度、消融效应)被当作因果支持,却未说明使其成立的前提。两名审稿人对n=30篇论文的复核重现了主发现:识别假设章节缺失、验证指标替代识别现象普遍,但具体维度计数受编码规则敏感影响。本文提出披露规范:声明是否为因果主张,命名识别策略,列出所有假设,强调至少一个核心假设,并解释若其不成立结论如何变化。验证不能替代识别。

原文摘要 · Abstract (English)

Mechanistic interpretability papers increasingly use causal vocabulary: circuits, mediators, causal abstraction, monosemanticity. Such claims require explicit identification assumptions. A purposive audit of 10 papers across four methodological strands finds no dedicated identification-assumptions section and a recurring pattern: validation metrics such as faithfulness, completeness, monosemanticity, alignment, or ablation effects are reported as causal support without stating the assumptions that make them identifying. A two-human-coder audit on $n=30$ reproduces the direction of the main finding: dedicated identification sections are absent, and validation-metric substitution is common, though exact Dim B/D counts are coding-rule sensitive. The paper proposes a disclosure norm: state whether the claim is causal, name the identification strategy, enumerate assumptions, stress at least one, and explain how conclusions shift if assumptions fail. Validation is not identification.

可解释性因果推断论文规范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。