提出可量化评估解释质量的新框架,无需真实标签即可训练出因果解释。
Learning Quantifiable Visual Explanations Without Ground-Truth

- 基于连续输入扰动构建解释质量度量,兼顾充分性与必要性。
- 新方法在多个指标上优于现有XAI技术,且不降低模型性能。
- 适用于任何黑盒模型,生成可信赖的因果解释,适合可信AI研究者。
可解释人工智能(XAI)对验证和负责任使用现代深度学习模型日益重要,但因缺乏可靠的真值进行对比而难以评估。本文提出一种基于连续输入扰动的可量化框架,用于衡量XAI方法的质量,该框架形式化地考虑了解释信息对模型决策的充分性与必要性。我们展示了该度量在多种情况下比现有指标更符合人类对解释质量的直觉。为进一步利用该度量,我们提出一种新型XAI方法:通过将该度量的可微分近似作为监督信号,对模型进行微调。结果是一个可附加于任意黑盒模型之上的适配器模块,能输出模型决策过程的因果解释,且不损害模型性能。实验表明,该方法生成的解释在多项可量化指标上优于现有XAI技术。
原文摘要 · Abstract (English)
Explainable AI (XAI) techniques are increasingly important for the validation and responsible use of modern deep learning models, but are difficult to evaluate due to the lack of good ground-truth to compare against. We propose a framework that serves as a quantifiable metric for the quality of XAI methods, based on continuous input perturbation. Our metric formally considers the sufficiency and necessity of the attributed information to the model's decision-making, and we illustrate a range of cases where it aligns better with human intuitions of explanation quality than do existing metrics. To exploit the properties of this metric, we also propose a novel XAI method, considering the case where we fine-tune a model using a differentiable approximation of the metric as a supervision signal. The result is an adapter module that can be trained on top of any black-box model to output causal explanations of the model's decision process, without degrading model performance. We show that the explanations generated by this method outperform those of competing XAI techniques according to a number of quantifiable metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。