用高阶概念解释视觉模型行为,揭示关键影响因素。
Concept-Based Abductive and Contrastive Explanations for Behaviors of Vision Models

- 基于概念擦除构建因果解释,找出影响预测的最小概念集。
- 可解释单张图及多张图共现的行为模式,支持用户自定义行为类型。
- 兼顾可读性与严谨性,适合需要透明化决策过程的研究者。
概念基解释为以人类可理解的概念解释深度神经网络预测提供了新路径。然而现有方法或未建立概念与预测间的因果关系,或表达能力有限,仅能处理单一概念的因果推理。与此同时,形式化归纳与对比解释虽能计算出导致模型结果的最小输入特征集,但仅关注像素等低级特征。本文融合两条研究路线,提出概念基归纳与对比解释,捕捉对模型输出具有因果关联的最小高阶概念集合。我们设计了一族算法,在使用概念擦除建立因果关系的同时枚举所有最小解释。通过合理聚合这些解释,不仅能理解单张图像的预测,还可分析模型在用户指定共性行为下的表现。我们在多个模型、数据集和行为上评估该方法,验证了其生成有益且易懂解释的有效性。
原文摘要 · Abstract (English)
*Concept-based explanations* offer a promising approach for explaining the predictions of deep neural networks in terms of high-level, human-understandable concepts. However, existing methods either do not establish a causal connection between the concepts and model predictions or are limited in expressivity and only able to infer causal explanations involving single concepts. At the same time, the parallel line of work on *formal abductive and contrastive explanations* computes the minimal set of input features causally relevant for model outcomes but only considers low-level features such as pixels. Merging these two threads, in this work, we propose the notion of *concept-based abductive and contrastive explanations* that capture the minimal sets of high-level concepts causally relevant for model outcomes. We then present a family of algorithms that enumerate all minimal explanations while using *concept erasure* procedures to establish causal relationships. By appropriately aggregating such explanations, we are not only able to understand model predictions on individual images but also on collections of images where the model exhibits a user-specified, common *behavior*. We evaluate our approach on multiple models, datasets, and behaviors, and demonstrate its effectiveness in computing helpful, user-friendly explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。