arXiv:2409.10329cs.CVcs.AI2024-09被引 6

通过信息解耦技术,让图像分类模型的决策变得可解释。

InfoDisent: Explainability of Image Classification Models by Information Disentanglement

  • 基于信息瓶颈原理,分解模型最后一层的特征为可解释的原子概念。
  • 在ViT和卷积网络上验证,能有效识别图像中的原型部件。
  • 适用于新领域(如ImageNet),适合需要模型可解释性的研究者。

本文提出InfoDisent,一种基于信息瓶颈原理的混合可解释性方法。该方法能将任意预训练模型最后层的信息解耦为原子概念,可解释为原型部件。该方法结合了后处理方法的灵活性与自解释神经网络(如ProtoPNets)的概念建模能力。我们在多种数据集上,使用ViTs和卷积网络等现代骨干网络,通过计算实验和用户研究验证了其有效性。值得注意的是,InfoDisent将原型部件方法推广至新领域(ImageNet),展现出良好的泛化能力。

原文摘要 · Abstract (English)

In this work, we introduce InfoDisent, a hybrid approach to explainability based on the information bottleneck principle. InfoDisent enables the disentanglement of information in the final layer of any pretrained model into atomic concepts, which can be interpreted as prototypical parts. This approach merges the flexibility of post-hoc methods with the concept-level modeling capabilities of self-explainable neural networks, such as ProtoPNets. We demonstrate the effectiveness of InfoDisent through computational experiments and user studies across various datasets using modern backbones such as ViTs and convolutional networks. Notably, InfoDisent generalizes the prototypical parts approach to novel domains (ImageNet).

可解释性图像分类信息瓶颈原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。