提出自动解析视觉概念的解释链,提升模型决策可解释性。
CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification
- 自动构建视觉概念描述,生成全局解释数据集
- 设计去耦合机制分离多义概念,用熵量化解释不确定性
- 支持语言化局部解释,适合需要透明决策的AI应用
可解释性是深度视觉模型(DVMs)广泛部署的关键。基于概念的后处理解释方法能提供全局与局部洞察,但现有方法难以自动构建准确、充分的语言化全局概念解释和局部电路描述。尤其,语义视觉概念(VCs)的内在多义性严重阻碍了解释能力,却常被低估。本文提出解释链(CoE)方法:首先自动解码并描述视觉概念,构建全局概念解释数据集;其次设计概念多义性解耦与过滤机制,区分最相关的概念原子;再引入概念多义性熵(CPE)度量模型可解释性,将确定性概念建模为不确定的概念原子分布;最终通过追踪概念电路,实现对DVM决策过程的自动语言化局部解释。GPT-4o与人工实验验证其有效性,解释评分平均绝对提升36%。
原文摘要 · Abstract (English)
Explainability is a critical factor influencing the wide deployment of deep vision models (DVMs). Concept-based post-hoc explanation methods can provide both global and local insights into model decisions. However, current methods in this field face challenges in that they are inflexible to automatically construct accurate and sufficient linguistic explanations for global concepts and local circuits. Particularly, the intrinsic polysemanticity in semantic Visual Concepts (VCs) impedes the interpretability of concepts and DVMs, which is underestimated severely. In this paper, we propose a Chain-of-Explanation (CoE) approach to address these issues. Specifically, CoE automates the decoding and description of VCs to construct global concept explanation datasets. Further, to alleviate the effect of polysemanticity on model explainability, we design a concept polysemanticity disentanglement and filtering mechanism to distinguish the most contextually relevant concept atoms. Besides, a Concept Polysemanticity Entropy (CPE), as a measure of model interpretability, is formulated to quantify the degree of concept uncertainty. The modeling of deterministic concepts is upgraded to uncertain concept atom distributions. Finally, CoE automatically enables linguistic local explanations of the decision-making process of DVMs by tracing the concept circuit. GPT-4o and human-based experiments demonstrate the effectiveness of CPE and the superiority of CoE, achieving an average absolute improvement of 36% in terms of explainability scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。