arXiv:2411.18645cs.CVcs.LG2024-11中稿 · IEEE WACV2026被引 2

让图像分类模型的决策过程可解释,看清它用了哪些概念做判断。

Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings

  • 通过双向交互模块,让概念与输入嵌入相互影响,实现多层次可解释性。
  • 能量化每个概念对预测的贡献,并定位其在图像中的具体位置。
  • 适合关注AI决策透明度的研究者和需要可解释模型的开发者。

内省可解释性旨在通过可扩展、自动化的方法揭示AI系统的内部机制。尽管大语言模型已有大量研究,但针对大规模图像任务的内省可解释性仍较少,主要集中在架构与功能层面以可视化学习到的概念。本文首次提出一个支持内省可解释性和多层级分析的大规模图像分类框架。具体而言,引入双向概念与输入嵌入交互模块(Bi-ICE),实现计算、算法与实现层面的可解释性。该模块通过基于人类可理解概念生成预测,量化其贡献并定位其在输入中的位置,提升透明度。实验展示了图像分类中概念贡献的度量与定位能力。方法通过展示概念学习过程及其收敛性,突出算法级可解释性。

原文摘要 · Abstract (English)

Inner interpretability is a promising field aiming to uncover the internal mechanisms of AI systems through scalable, automated methods. While significant research has been conducted on large language models, limited attention has been paid to applying inner interpretability to large-scale image tasks, focusing primarily on architectural and functional levels to visualize learned concepts. In this paper, we first present a conceptual framework that supports inner interpretability and multilevel analysis for large-scale image classification tasks. Specifically, we introduce the Bi-directional Interaction between Concept and Input Embeddings (Bi-ICE) module, which facilitates interpretability across the computational, algorithmic, and implementation levels. This module enhances transparency by generating predictions based on human-understandable concepts, quantifying their contributions, and localizing them within the inputs. Finally, we showcase enhanced transparency in image classification, measuring concept contributions, and pinpointing their locations within the inputs. Our approach highlights algorithmic interpretability by demonstrating the process of concept learning and its convergence.

可解释性图像分类概念学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。