arXiv:2605.05480cs.LGcs.AI2026-05

融合梯度与联盟方法,实现可证明的归因精度与采样稳定性。

GRALIS: Fusing Coalition and Gradient Attribution with Closed-Form Conservation Error and Finite-Sample Guarantees

  • 基于瑞斯表示定理,统一梯度与联盟归因机制
  • 闭式完整性误差保证,且采样误差收敛于1/√m + 1/k²
  • 适用于需要可验证解释的医学影像等高可靠性场景

现有后验可解释性方法(如GradCAM、SHAP、LIME、积分梯度)理论基础各异,难以统一比较。近期研究发现其联盟类与梯度类方法在不同忠实性指标上互补,仅靠方法选择无法根本解决。本文提出GRALIS,将谢帕利联盟权重与局部核,以及连续积分梯度路径融合为单一估计器,并提供两种独有保证:闭式完整性缺陷(阶数为d的交互项仅被分配真实值的1/d,当使用谢帕利权重与双线性函数时);以及实际自归一化比值的有限样本界,为O(1/√m) + O(1/k²)。该融合基于表示论结果:所有可加、线性、连续的归因泛函均可通过瑞斯表示定理唯一表示,逐特征独立证明而非跨特征共享形式。该类包含SHAP、IG和LIME,但不包括非线性泛函如标准GradCAM或注意力图。七条定理进一步建立了与谢帕利交互值的精确对应、与霍夫丁/Sobol分解的仿射对应关系,以及最小方差多尺度扩展。初步实验展示在乳腺组织病理图像上的应用;扩展验证见附录论文(Fanale, 2026)。

原文摘要 · Abstract (English)

The main post-hoc XAI methods for deep networks -- GradCAM, SHAP, LIME, Integrated Gradients -- originate from heterogeneous theoretical foundations and are not naturally comparable within a single representation. A recent benchmark also finds their coalition-based members (GradCAM, KernelSHAP, LIME) and gradient-based members (Integrated Gradients and variants) empirically complementary, each outperforming the other on different faithfulness metrics, with method selection as the only proposed remedy (Gevaert et al., 2022). This work presents GRALIS (Gradient-Riesz Averaged Locally-Integrated Shapley), which fuses these two mechanisms -- a Shapley coalition weight and locality kernel, and a continuous Integrated-Gradients-style conditioned path -- into a single estimator, and equips it with two certified guarantees neither mechanism supplies alone: an exact, closed-form completeness deficit (an order-d interaction is attributed at a factor 1/d of its true value under Shapley weights and a multilinear F) and a finite-sample bound, O(1/sqrt(m)) + O(1/k^2), for the actual self-normalized ratio the algorithm returns. This fusion is underpinned by a representation-theoretic result: every additive, linear, continuous attribution functional admits a unique canonical representation via the Riesz Representation Theorem, proved componentwise (feature by feature) rather than as one form shared across features or methods. This class includes SHAP, IG and LIME, but not nonlinear functionals such as standard GradCAM or attention maps. Seven theorems further establish an exact correspondence with Shapley Interaction Values, affine-regime correspondences with the Hoeffding/Sobol decomposition, and a minimum-variance multi-scale extension. A preliminary experimental illustration on breast histology imaging is included; extended validation is in a companion paper (Fanale, 2026).

可解释AI归因方法理论保证医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。