arXiv:2410.12439cs.LG2024-10被引 1

统一解释形式,让模型说明更直观可信。

Beyond Attribution: Unified Concept-Level Explanations

  • 用预训练模型扰动统一扩展现有方法,生成概念级解释。
  • 支持归因、充分条件和反事实三种解释形式,效果优于当前最佳。
  • 适用于文本、图像、多模态模型,适合需要多样解释的用户。

随着对可解释性的需求增加,模型无关解释与基于概念的方法结合日益重要:前者适用多种模型架构,后者使解释更忠实且易于理解。然而,现有基于概念的模型无关解释方法范围有限,主要聚焦于归因式解释,忽略充分条件和反事实等多样化形式,限制了实用性。为此,我们提出通用框架UnCLE,将现有局部模型无关技术扩展为统一的概念级解释。核心思想是利用大规模预训练模型扰动,统一扩展现有方法以生成概念级解释。我们已实现UnCLE在三种形式上的应用:归因、充分条件与反事实,并应用于主流文本、图像及多模态模型。评估表明,UnCLE提供的解释比现有最佳方法更忠实,且形式更丰富,满足不同用户需求。

原文摘要 · Abstract (English)

There is an increasing need to integrate model-agnostic explanation techniques with concept-based approaches, as the former can explain models across different architectures while the latter makes explanations more faithful and understandable to end-users. However, existing concept-based model-agnostic explanation methods are limited in scope, mainly focusing on attribution-based explanations while neglecting diverse forms like sufficient conditions and counterfactuals, thus narrowing their utility. To bridge this gap, we propose a general framework UnCLE to elevate existing local model-agnostic techniques to provide concept-based explanations. Our key insight is that we can uniformly extend existing local model-agnostic methods to provide unified concept-based explanations with large pre-trained model perturbation. We have instantiated UnCLE to provide concept-based explanations in three forms: attributions, sufficient conditions, and counterfactuals, and applied it to popular text, image, and multimodal models. Our evaluation results demonstrate that UnCLE provides explanations more faithful than state-of-the-art concept-based explanation methods, and provides richer explanation forms that satisfy various user needs.

可解释性概念解释多模态统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。