用元博弈框架量化解释方法的相互影响,揭示深层特征关系。
Attributions All the Way Down? The Metagame of Interpretability

- 将解释方法视为合作博弈,计算特征对解释的影响力。
- 在语言、视觉语言和图文生成模型中发现显著的特征交互模式。
- 适合研究解释可信赖性与多模态模型内部机制的学者。
我们提出元博弈框架,用于量化模型解释的二阶交互效应。对于任意解释模型 $f$ 的一阶归因 $ϕ(f)$,通过将归因方法本身视为合作博弈并计算其谢林值,衡量特征 $j$ 对特征 $i$ 归因的定向影响,记为元归因 $φ_{j o i}(f)$。理论上,我们证明归因可逐层分解为元归因,并将其作为现有交互指数的方向性扩展。实证上,元博弈在多种可解释性任务中展现价值:(i) 量化指令微调语言模型中的标记交互;(ii) 解释视觉-语言编码器中的跨模态相似性;(iii) 解读多模态扩散变换器中的文生图概念。
原文摘要 · Abstract (English)
We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution $ϕ(f)$ explaining a model $f$, we measure the directional influence of feature $j$ on the attribution of feature $i$, denoted as meta-attribution $φ_{j \to i}(f)$, by treating the attribution method itself as a cooperative game and computing its Shapley value. Theoretically, we prove that attributions hierarchically decompose into meta-attributions, and establish these as directional extensions of existing interaction indices. Empirically, we demonstrate that the metagame delivers insights across diverse interpretability applications: (i) quantifying token interactions in instruction-tuned language models, (ii) explaining cross-modal similarity in vision-language encoders, and (iii) interpreting text-to-image concepts in multimodal diffusion transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。