改进反事实解释方法,让AI决策理由更可信且易懂。
Info-CELS: Informative Saliency Map Guided Counterfactual Explanation
- 用可解释性热力图引导生成反事实解释
- 在保持接近原样本的前提下,提升解释有效性
- 适合需要高可信度解释的AI应用开发者
随着对可解释机器学习需求的增长,人类参与提供有意义的模型决策解释变得愈发重要,这有助于建立AI系统的信任与透明。为此,可解释人工智能(XAI)领域应运而生。近期提出的CELS模型首次利用学习到的显著性图,既直观说明时序分类器决策原因,又生成后验反事实解释。然而,该模型在保证高近似性和稀疏性的同时,存在有效性不足的问题。本文提出改进方法Info-CELS,通过去除掩码归一化,使解释更具信息量和有效性。在多个领域数据集上的大量实验表明,该方法优于原始CELS,在有效性和解释信息量上均有提升。
原文摘要 · Abstract (English)
As the demand for interpretable machine learning approaches continues to grow, there is an increasing necessity for human involvement in providing informative explanations for model decisions. This is necessary for building trust and transparency in AI-based systems, leading to the emergence of the Explainable Artificial Intelligence (XAI) field. Recently, a novel counterfactual explanation model, CELS, has been introduced. CELS learns a saliency map for the interest of an instance and generates a counterfactual explanation guided by the learned saliency map. While CELS represents the first attempt to exploit learned saliency maps not only to provide intuitive explanations for the reason behind the decision made by the time series classifier but also to explore post hoc counterfactual explanations, it exhibits limitations in terms of high validity for the sake of ensuring high proximity and sparsity. In this paper, we present an enhanced approach that builds upon CELS. While the original model achieved promising results in terms of sparsity and proximity, it faced limitations in validity. Our proposed method addresses this limitation by removing mask normalization to provide more informative and valid counterfactual explanations. Through extensive experimentation on datasets from various domains, we demonstrate that our approach outperforms the CELS model, achieving higher validity and producing more informative explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。