提出可忠实追踪的通用概念解释方法,提升模型可解释性。
FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
- 基于模型内在机制提取跨类别共享概念
- 概念贡献与可视化可精准追溯到每一层
- 新评估指标验证解释一致性,适合可信AI研究者
深度网络在众多任务中表现卓越,但获得全局概念级理解仍是关键挑战。现有后处理概念解释方法常不忠实于模型,且对概念设定强假设,如类别专属、空间范围小或符合人类预期。本文强调解释的忠实性,提出一种具备模型内生机制的概念解释新方法。概念跨类别共享,可从任意层追溯其对输出logit的贡献及输入可视化。同时利用基础模型提出新的概念一致性度量C²-Score,用于评估概念解释方法。实验表明,相比以往方法,本方法在概念一致性上更优,用户认为解释更易懂,同时保持ImageNet上的竞争力性能。
原文摘要 · Abstract (English)
Deep networks have shown remarkable performance across a wide range of tasks, yet getting a global concept-level understanding of how they function remains a key challenge. Many post-hoc concept-based approaches have been introduced to understand their workings, yet they are not always faithful to the model. Further, they make restrictive assumptions on the concepts a model learns, such as class-specificity, small spatial extent, or alignment to human expectations. In this work, we put emphasis on the faithfulness of such concept-based explanations and propose a new model with model-inherent mechanistic concept-explanations. Our concepts are shared across classes and, from any layer, their contribution to the logit and their input-visualization can be faithfully traced. We also leverage foundation models to propose a new concept-consistency metric, C$^2$-Score, that can be used to evaluate concept-based methods. We show that, compared to prior work, our concepts are quantitatively more consistent and users find our concepts to be more interpretable, all while retaining competitive ImageNet performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。