arXiv:2508.12947stat.MLcs.CE2025-08被引 3

提出改进的配对采样法,更准更快计算机器学习解释中的贡献值。

Shapley Values: Paired-Sampling Approximations

  • 用配对采样替代传统采样,提升估计精度。
  • 在二阶交互下可得精确解,且具可加性恢复特性。
  • 适合需要高精度解释的模型调试与可信分析场景。

Shapley值源于合作博弈论,现已成为解释机器学习预测的重要工具。依据公平性公理,每个输入特征获得与其对输出贡献相匹配的信用值,用于解释预测结果。计算这些信用值的主要限制是计算复杂度。现有两种主流采样近似方法:采样KernelSHAP和采样PermutationSHAP。本文首次提供这两种方法的渐近正态性结果。进一步证明,在交互项最高为二阶的情况下,配对采样方法可获得精确结果。此外,配对采样PermutationSHAP具备可加性恢复性质,而其核版本则不具备。

原文摘要 · Abstract (English)

Originally introduced in cooperative game theory, Shapley values have become a very popular tool to explain machine learning predictions. Based on Shapley's fairness axioms, every input (feature component) gets a credit how it contributes to an output (prediction). These credits are then used to explain the prediction. The only limitation in computing the Shapley values (credits) for many different predictions is of computational nature. There are two popular sampling approximations, sampling KernelSHAP and sampling PermutationSHAP. Our first novel contributions are asymptotic normality results for these sampling approximations. Next, we show that the paired-sampling approaches provide exact results in case of interactions being of maximal order two. Furthermore, the paired-sampling PermutationSHAP possesses the additive recovery property, whereas its kernel counterpart does not.

模型解释Shapley值采样优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。