arXiv:2602.16698cs.LG2026-02被引 8

用因果框架提升大模型可解释性结论的泛化能力

Causality is Key for Interpretability Claims to Generalise

  • 基于贝叶斯因果层级,区分观察、干预与反事实三类证据
  • 指出当前可解释性研究多依赖不可验证的反事实假设
  • 提出诊断框架,助研究者匹配方法与可信结论

大语言模型可解释性研究虽取得重要进展,但普遍存在结论无法泛化、因果推断超越证据等问题。本文认为,因果推断明确了从模型激活值到稳定高层结构的有效映射所需的数据或假设,以及可支持的推论类型。具体而言,佩尔的因果层级阐明了可解释性研究可成立的边界:观测仅能建立行为与内部组件间的关联;干预(如消融或激活修补)可支持编辑对特定提示集下行为指标(如平均词元概率变化)的影响;而反事实命题——即同一提示在未观测干预下的输出——在缺乏受控监督的情况下基本不可验证。我们展示因果表示学习(CRL)如何实现这一层级,明确哪些变量可从激活中恢复及对应假设。由此提出诊断框架,帮助研究者选择与声明相匹配的方法与评估,确保结论具有泛化性。

原文摘要 · Abstract (English)

Interpretability research on large language models (LLMs) has yielded important insights into model behaviour, yet recurring pitfalls persist: findings that do not generalise, and causal interpretations that outrun the evidence. Our position is that causal inference specifies what constitutes a valid mapping from model activations to invariant high-level structures, the data or assumptions needed to achieve it, and the inferences it can support. Specifically, Pearl's causal hierarchy clarifies what an interpretability study can justify. Observations establish associations between model behaviour and internal components. Interventions (e.g., ablations or activation patching) support claims how these edits affect a behavioural metric (e.g., average change in token probabilities) over a set of prompts. However, counterfactual claims -- i.e., asking what the model output would have been for the same prompt under an unobserved intervention -- remain largely unverifiable without controlled supervision. We show how causal representation learning (CRL) operationalises this hierarchy, specifying which variables are recoverable from activations and under what assumptions. Together, these motivate a diagnostic framework that helps practitioners select methods and evaluations matching claims to evidence such that findings generalise.

可解释性因果推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。