用博弈论框架统一解释模型决策的归因方法。
Playing the network backward: A Game Theoretic Attribution Framework

- 将归因视为网络上的双人博弈,统一梯度与LRP等方法
- 新方法在ViT-B/16上超越现有Transformer归因指标
- 适合关注模型可解释性与鲁棒性的研究者
归因方法用于解释哪些输入特征驱动了模型预测,是模型调试和机制可解释性的核心。然而,现有的反向归因方法(如梯度、LRP及Transformer特有规则)缺乏共享的理论框架来比较其底层计算逻辑。本文通过将反向归因重新建模为扩展网络图上的双人博弈,构建了统一框架。梯度与完整的alpha-beta-LRP族均成为特定均衡下博弈轨迹的积分形式,归因图谱变为轨迹分布的投影而非核心对象。期望的解释特性(如定位精度、对输入噪声的鲁棒性、稳定注意力路由)可转化为博弈论概念(如策略正则化、风险规避、扩展动作集),并直接映射为已有反向规则的新变体。在ViT-B/16上,一种选定的alpha-beta-LRP改进版本在所有定位指标上均优于先前的Transformer专用方法。
原文摘要 · Abstract (English)
Attribution methods explain which input features drive a model's prediction, making them central to model debugging and mechanistic interpretability. Yet backward attribution methods, including gradients, LRP, and transformer-specific rules, lack a shared framework in which to compare the underlying backward calculations. We introduce such a framework by recasting backward attribution as a two-player game on an extended network graph, building on Gaubert and Vlassopoulos' ReLU Net Game. Gradients and the full alpha-beta-LRP family arise as integrals over game trajectories under specific equilibria, so attribution maps become projections of trajectory distributions rather than the primary object. Desired explanation properties, such as localisation focus, robustness to input noise, or stable attention routing, can be specified as game-theoretic concepts, including policy regularization, risk aversion, and extended action sets, and translate directly into novel adaptations of the well-known backward rules. On ViT-B/16, one such selected adaptation of alpha-beta-LRP outperforms prior transformer-specific backward methods across all considered localisation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。