arXiv:2410.17832cs.CV2024-10被引 3

用图文嵌入空间自动生成概念描述,提升神经网络可解释性。

Exploiting Text-Image Latent Spaces for the Description of Visual Concepts

  • 通过提取关键感受野而非整图编码,映射到图文嵌入空间。
  • 在有无标签情况下均生成准确的概念描述,提升解释效率。
  • 适合需要理解模型决策依据的研究者与开发者使用。

概念激活向量(CAVs)通过将人类可理解的概念与神经网络内部特征提取过程关联,揭示其决策机制。然而,当发现新的CAV集合时,仍需人工将其转化为可读描述。对于图像模型,通常通过可视化最相关图像来辅助理解,但概念定义仍依赖人工判断。本文提出一种新方法,通过将代表CAV的最相关图像映射至图文嵌入空间,自动计算这些图像的联合文本描述,从而辅助解释。我们采用最相关的感受野而非完整图像进行编码,提升表征精度。在多个有/无标签的实验中,该方法均能生成准确且有意义的概念描述,显著降低概念解释难度。

原文摘要 · Abstract (English)

Concept Activation Vectors (CAVs) offer insights into neural network decision-making by linking human friendly concepts to the model's internal feature extraction process. However, when a new set of CAVs is discovered, they must still be translated into a human understandable description. For image-based neural networks, this is typically done by visualizing the most relevant images of a CAV, while the determination of the concept is left to humans. In this work, we introduce an approach to aid the interpretation of newly discovered concept sets by suggesting textual descriptions for each CAV. This is done by mapping the most relevant images representing a CAV into a text-image embedding where a joint description of these relevant images can be computed. We propose utilizing the most relevant receptive fields instead of full images encoded. We demonstrate the capabilities of this approach in multiple experiments with and without given CAV labels, showing that the proposed approach provides accurate descriptions for the CAVs and reduces the challenge of concept interpretation.

可解释性图文嵌入CAV

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。