提出新方法OddSHAP,让特征重要性计算更准更快。
An Odd Estimator for Shapley Values
- 发现谢泼德值仅依赖集合函数的奇部,据此设计新估计器。
- 在多个数据集上,采样量更大时准确率领先现有方法。
- 适合需要高精度解释的机器学习模型分析场景。
谢泼德值是机器学习中广泛使用的归因框架,涵盖特征重要性、数据估值和因果推断。然而其精确计算通常不可行,需依赖高效近似方法。尽管最有效且流行的方法利用成对采样策略降低估计误差,但该机制的理论基础一直不清晰。本文提供了一个简洁而根本的解释:我们证明谢泼德值仅依赖于集合函数的奇部,而成对采样通过正交化回归目标,滤除了无关的偶部成分。基于此,我们提出OddSHAP,一种新的一致估计器,仅在奇子空间进行多项式回归。通过傅里叶基分离该子空间,并使用代理模型识别高影响交互,OddSHAP克服了高阶近似中的组合爆炸问题。在广泛基准测试中,我们发现OddSHAP在较大采样预算下实现了最先进的估计精度。
原文摘要 · Abstract (English)
The Shapley value is a ubiquitous framework for attribution in machine learning, encompassing feature importance, data valuation, and causal inference. However, its exact computation is generally intractable, necessitating efficient approximation methods. While the most effective and popular estimators leverage the paired sampling heuristic to reduce estimation error, the theoretical mechanism driving this improvement has remained opaque. In this work, we provide an elegant and fundamental justification for paired sampling: we prove that the Shapley value depends exclusively on the odd component of the set function, and that paired sampling orthogonalizes the regression objective to filter out the irrelevant even component. Leveraging this insight, we propose OddSHAP, a novel consistent estimator that performs polynomial regression solely on the odd subspace. By utilizing the Fourier basis to isolate this subspace and employing a proxy model to identify high-impact interactions, OddSHAP overcomes the combinatorial explosion of higher-order approximations. Through an extensive benchmark, we find that OddSHAP achieves state-of-the-art estimation accuracy at larger sampling budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。