arXiv:2602.11448cs.LGcs.CV2026-02被引 1

让图像分类模型的解释更符合语义层级,提升可解释性。

Hierarchical Concept Embedding & Pursuit for Interpretable Image Classification

  • 构建语义概念的层次嵌入结构,用分层稀疏编码恢复图像中的概念。
  • 在真实数据集上,概念精确率和召回率均优于现有方法。
  • 特别适合小样本场景,兼具高准确率与可解释性,适合医疗等关键领域。

可解释性设计的模型在计算机视觉中日益受到关注,因其能提供预测的可信解释。在图像分类任务中,这类模型通常从图像中恢复人类可理解的概念并用于分类。现有的稀疏概念恢复方法利用视觉-语言模型的潜在空间,将图像嵌入表示为概念嵌入的稀疏组合。然而,这些方法忽略了语义概念的层级结构,可能导致预测正确但解释与层级不一致。本文提出分层概念嵌入与追踪框架(HCEP),在潜在空间中引入概念的层次结构,并进行分层稀疏编码以恢复图像中存在的概念。给定一个语义概念的层次结构,我们提出一种几何构造方法来生成对应的嵌入层次。假设真实概念构成层次中的根路径,我们推导出其在嵌入空间中可恢复的充分条件。实验表明,分层稀疏编码能可靠恢复分层概念嵌入,而标准稀疏编码则失败。在真实世界数据集上的实验显示,相较于现有方法,HCEP在保持竞争性分类准确率的同时,显著提升了概念精度与召回率。尤其在用于概念估计和分类器训练的样本数量有限时,HCEP实现了更高的分类准确率与概念恢复效果。结果表明,将层次结构融入稀疏概念恢复,能构建更忠实、更具可解释性的图像分类模型。

原文摘要 · Abstract (English)

Interpretable-by-design models are gaining traction in computer vision because they provide faithful explanations for their predictions. In image classification, these models typically recover human-interpretable concepts from an image and use them for classification. Sparse concept recovery methods leverage the latent space of vision-language models to represent image embeddings as sparse combinations of concept embeddings. However, by ignoring the hierarchical structure of semantic concepts, these methods may produce correct predictions with explanations that are inconsistent with the hierarchy. In this work, we propose Hierarchical Concept Embedding & Pursuit (HCEP), a framework that induces a hierarchy of concept embeddings in the latent space and performs hierarchical sparse coding to recover the concepts present in an image. Given a hierarchy of semantic concepts, we introduce a geometric construction for the corresponding hierarchy of embeddings. Under the assumption that the true concepts form a rooted path in the hierarchy, we derive sufficient conditions for their recovery in the embedding space. We further show that hierarchical sparse coding reliably recovers hierarchical concept embeddings, whereas standard sparse coding fails. Experiments on real-world datasets show that HCEP improves concept precision and recall compared to existing methods while maintaining competitive classification accuracy. Moreover, when the number of samples available for concept estimation and classifier training is limited, HCEP achieves superior classification accuracy and concept recovery. Our results demonstrate that incorporating hierarchical structure into sparse concept recovery leads to more faithful and interpretable image classification models.

可解释性层次结构稀疏编码图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。