提出新采样方法,高效找出最重要的k个特征。
Antithetic Sampling for Top-k Shapley Identification
- 利用相关性设计新采样策略,提升重要特征识别效率
- 实验证明全特征估计与重点特征识别效果不互通
- 适合关注关键特征的可解释性场景
加法型特征解释主要依赖博弈论中的沙普利值,因其公理唯一性而广受欢迎。但其计算复杂度严重制约实际应用。现有方法多聚焦于所有特征沙普利值的均匀近似,对无关紧要特征也消耗大量样本。相比之下,识别前k个最重要特征已具足够洞察力,并可结合多臂老虎机算法优化。本文提出可比边际贡献采样(CMCS),利用相关观测设计新型采样方案,解决顶k特征识别问题。实验表明,全特征近似效果与顶k识别效果并不一致,两者不可互换。该方法在多个基准上优于现有方法。
原文摘要 · Abstract (English)
Additive feature explanations rely primarily on game-theoretic notions such as the Shapley value by viewing features as cooperating players. The Shapley value's popularity in and outside of explainable AI stems from its axiomatic uniqueness. However, its computational complexity severely limits practicability. Most works investigate the uniform approximation of all features' Shapley values, needlessly consuming samples for insignificant features. In contrast, identifying the $k$ most important features can already be sufficiently insightful and yields the potential to leverage algorithmic opportunities connected to the field of multi-armed bandits. We propose Comparable Marginal Contributions Sampling (CMCS), a method for the top-$k$ identification problem utilizing a new sampling scheme taking advantage of correlated observations. We conduct experiments to showcase the efficacy of our method in compared to competitive baselines. Our empirical findings reveal that estimation quality for the approximate-all problem does not necessarily transfer to top-$k$ identification and vice versa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。