arXiv:2409.01610cs.CVcs.AI2024-09被引 1

通过分解模型路径,揭示图像模型内部语义运作机制。

Decompose the model: Mechanistic interpretability in image models with Generalized Integrated Gradients (GIG)

  • 用点特征向量与有效感受野分解嵌入为可解释概念向量
  • 基于广义积分梯度计算概念间相关性,实现全数据集分析
  • 适用于想理解图像模型内部工作原理的研究者

在自然语言模型的可解释人工智能(XAI)领域,从个体决策的局部解释发展到包含高层概念的全局解释,已为机制可解释性奠定基础,旨在解码模型的具体操作。然而这一范式在图像模型中尚未充分探索,现有方法多集中于类别特定的解释。本文提出一种新方法,系统追踪从输入经所有中间层到最终输出在整个数据集中的完整路径。我们利用点特征向量(PFVs)和有效感受野(ERFs)将模型嵌入分解为可解释的概念向量,并通过广义积分梯度(GIG)计算概念向量间的相关性,实现对模型行为的全面、数据集级别的分析。我们在定性和定量评估中验证了概念提取与概念归因的有效性。该方法深化了对图像模型中语义意义的理解,提供了模型运作机制的全景视图。

原文摘要 · Abstract (English)

In the field of eXplainable AI (XAI) in language models, the progression from local explanations of individual decisions to global explanations with high-level concepts has laid the groundwork for mechanistic interpretability, which aims to decode the exact operations. However, this paradigm has not been adequately explored in image models, where existing methods have primarily focused on class-specific interpretations. This paper introduces a novel approach to systematically trace the entire pathway from input through all intermediate layers to the final output within the whole dataset. We utilize Pointwise Feature Vectors (PFVs) and Effective Receptive Fields (ERFs) to decompose model embeddings into interpretable Concept Vectors. Then, we calculate the relevance between concept vectors with our Generalized Integrated Gradients (GIG), enabling a comprehensive, dataset-wide analysis of model behavior. We validate our method of concept extraction and concept attribution in both qualitative and quantitative evaluations. Our approach advances the understanding of semantic significance within image models, offering a holistic view of their operational mechanics.

图像模型可解释性机制解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。