提出高效准确的交互效应估算方法,解决高维模型解释难题。
Proxy-Based Approximation of Shapley and Banzhaf Interactions

- 用树模型代理+残差修正,兼顾速度与精度
- 理论证明可精确计算树集成的交互指数,避免指数级复杂度
- 在千维特征场景下仍保持低误差,适合大规模解释任务
Shapley和Banzhaf交互效应能捕捉现代机器学习应用中的复杂动态。然而现有估算方法在速度与精度间权衡不佳。为此,我们提出ProxySHAP,将基于树的代理模型的高采样效率与通过残差校正实现一致性的理论路径相结合。理论上,我们推导出干预型TreeSHAP的多项式时间推广,可对树集成精确计算交互指数,克服了先前方法中依赖树深度的指数复杂度。此外,我们形式化分析了残差调整策略,刻画了最大样本重用(MSR)在不使方差随交互规模指数增长的前提下纠正代理偏差的具体条件。大量基准测试表明,ProxySHAP在近似质量上达到新最优标准,包括上千特征的大规模应用。在小预算和大预算场景下均以最低误差超越前代最优方法ProxySPEX和KernelSHAP-IQ,且在下游可解释性任务中表现更优。
原文摘要 · Abstract (English)
Shapley and Banzhaf interactions capture the complex dynamics inherent in modern machine learning applications. However, current estimators for these higher-order interactions trade off between speed and accuracy. To overcome this limitation, we introduce ProxySHAP. ProxySHAP reconciles the high sample efficiency of tree-based proxy models with a principled path to consistency via residual correction. On a theoretical level, we derive a polynomial-time generalization of interventional TreeSHAP to compute exact interaction indices for tree ensembles, successfully bypassing exponential tree-depth dependencies in prior methods. Furthermore, we formally analyze the residual adjustment strategy, characterizing the specific conditions under which Maximum Sample Reuse (MSR) corrects proxy bias without its variance scaling exponentially with interaction size. Extensive benchmarking demonstrates that ProxySHAP sets a new state-of-the-art standard for approximation quality, including in large-scale applications with thousands of features. By achieving the lowest error in both small- and large-budget regimes, ProxySHAP significantly outperforms the prior best estimators ProxySPEX and KernelSHAP-IQ, while also delivering superior performance on downstream explainability tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。