提出新方法评估解释器稳定性,量化特征重要性排序的可靠性。
Measuring Explainer Stability via Attribution Separability

- 基于分布分析特征重要性排序的可分性,识别可靠排名的边界。
- 通过实验验证不同解释方法在数据集上的排序鲁棒性差异。
- 适合关注模型解释可信度的研究者与实践者使用。
特征重要性赋值方法(AMs)广泛用于解释黑箱模型,但多数方法因定义中的随机成分,会产生波动的赋值结果。本文提出一种基于分布的框架,用于捕捉赋值结果的稳定性。该方法能衡量排序向量中特征重要性的可分程度,并确定在何种条件下特征排序仍具可靠性。进一步地,该框架可用于比较不同方法在数据集上排序的稳健性。实验表明,该方法可有效评估解释器的稳定性,为现有评价体系提供互补标准。
原文摘要 · Abstract (English)
Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definition. In this paper, we propose a distribution-based framework to capture the stability of attribution scores. In particular, our approach allows to understand the degree of separability in the ranked attribution vector and obtain the largest index for which a feature ranking remains reliable. We further extend this framework to compare AMs based on the robustness of their rankings across a dataset. Through experiments, we demonstrate how to apply our method to evaluate explainer stability. Overall, our approach provides a complementary criterion for evaluating the stability of AMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。