提出新型稳定性认证方法,让解释结果更可靠且不依赖过度平滑模型。
Probabilistic Stability Guarantees for Feature Attributions
- 引入软稳定性概念,设计无需模型假设的高效认证算法
- 在视觉与语言任务上验证,比传统方法给出更优的稳定性保证
- 适用于各类归因方法,尤其适合追求解释可信度的研究者
稳定性保障已成为评估特征归因的合理方法,但现有认证方法依赖高度平滑的分类器,常导致过于保守的结论。为解决这一问题,我们提出软稳定性,并设计一种简单、模型无关、样本高效的稳定性认证算法(SCA),可为任意归因方法提供非平凡且可解释的保障。此外,我们发现适度平滑在准确率与稳定性之间取得更优权衡,避免了以往方法的极端妥协。通过布尔函数分析,我们推导出平滑下稳定性的新刻画。在视觉与语言任务上对SCA进行评估,证明软稳定性能有效衡量解释方法的鲁棒性。
原文摘要 · Abstract (English)
Stability guarantees have emerged as a principled way to evaluate feature attributions, but existing certification methods rely on heavily smoothed classifiers and often produce conservative guarantees. To address these limitations, we introduce soft stability and propose a simple, model-agnostic, sample-efficient stability certification algorithm (SCA) that yields non-trivial and interpretable guarantees for any attribution method. Moreover, we show that mild smoothing achieves a more favorable trade-off between accuracy and stability, avoiding the aggressive compromises made in prior certification methods. To explain this behavior, we use Boolean function analysis to derive a novel characterization of stability under smoothing. We evaluate SCA on vision and language tasks and demonstrate the effectiveness of soft stability in measuring the robustness of explanation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。