提出因果概念解释框架,让黑箱模型决策更透明可信。
A Framework for Causal Concept-based Model Explanations
- 基于概念干预的概率计算生成局部与全局解释
- 在CelebA数据集上验证了解释的可理解性与保真度
- 适合需要可解释性且关注因果关系的研究者
本文提出一种基于因果概念的后验可解释人工智能(XAI)概念框架,要求对非可解释模型的解释既易于理解又忠实于原模型。通过计算概念干预的充分性概率,生成局部与全局解释。以在CelebA数据集上训练的分类器为对象,展示了概念化词汇表的清晰表达能力,并隐含因果解读。通过强调框架的关键假设,强调解释生成与解释解读的上下文必须一致,从而保障解释的保真度。
原文摘要 · Abstract (English)
This work presents a conceptual framework for causal concept-based post-hoc Explainable Artificial Intelligence (XAI), based on the requirements that explanations for non-interpretable models should be understandable as well as faithful to the model being explained. Local and global explanations are generated by calculating the probability of sufficiency of concept interventions. Example explanations are presented, generated with a proof-of-concept model made to explain classifiers trained on the CelebA dataset. Understandability is demonstrated through a clear concept-based vocabulary, subject to an implicit causal interpretation. Fidelity is addressed by highlighting important framework assumptions, stressing that the context of explanation interpretation must align with the context of explanation generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。