用博弈论方法更全面评估特征重要性,提升模型可解释性。
Rigorous Feature Importance Scores based on Shapley Value and Banzhaf Index
- 引入谢帕利值和巴任指数,兼顾非弱归纳解释集的特征贡献。
- 新指标能量化特征排除对抗样本的有效性,提升可靠性。
- 适用于高风险场景下机器学习模型的可信解释需求。
基于博弈论的特征归因方法在可解释人工智能(XAI)领域广泛应用。近期工作采用基于逻辑的解释,特别是针对机器学习模型在高风险场景中的应用,提出严格的特征归因方法。通常这些方法使用弱归纳解释(WAXp)作为特征函数来分配特征重要性,但其缺点是忽略了非WAXp集合的贡献。实际上,非WAXp集合也可能包含重要信息,因为形式化解释(XPs)与对抗样本(AExs)之间存在关联。为此,本文利用谢帕利值(Shapley value)和巴任指数(Banzhaf index),提出两种新的特征重要性评分。在计算特征贡献时,考虑了非WAXp集合,并且新评分能够量化每个特征排除对抗样本的有效性。此外,论文还分析了所提评分的性质并研究其计算复杂度。
原文摘要 · Abstract (English)
Feature attribution methods based on game theory are ubiquitous in the field of eXplainable Artificial Intelligence (XAI). Recent works proposed rigorous feature attribution using logic-based explanations, specifically targeting high-stakes uses of machine learning (ML) models. Typically, such works exploit weak abductive explanation (WAXp) as the characteristic function to assign importance to features. However, one possible downside is that the contribution of non-WAXp sets is neglected. In fact, non-WAXp sets can also convey important information, because of the relationship between formal explanations (XPs) and adversarial examples (AExs). Accordingly, this paper leverages Shapley value and Banzhaf index to devise two novel feature importance scores. We take into account non-WAXp sets when computing feature contribution, and the novel scores quantify how effective each feature is at excluding AExs. Furthermore, the paper identifies properties and studies the computational complexity of the proposed scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。