arXiv:2601.12303cs.CV2026-01AAAI

让黑箱模型自动提取可解释的视觉概念,提升透明度与性能。

Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual Representations

  • 从预训练模型中分解出视觉概念,用多模态大模型筛选可靠概念。
  • 在11个分类任务上达到当前最佳准确率,接近端到端模型表现。
  • 适合需要模型可解释性的医疗、金融等高风险场景应用。

深度学习在图像识别中取得显著成功,但其内在不透明性限制了在关键领域的应用。基于概念的解释试图通过人类可理解的概念来阐明模型推理过程。然而,现有事后方法和事前概念瓶颈模型(CBMs)存在概念相关性不可靠、概念定义非可视化或耗时、以及模型或数据无关假设等问题。本文提出事后概念瓶颈模型通过表示分解(PCBM-ReD),一种将可解释性回溯至预训练黑箱模型的新流程。PCBM-ReD自动从预训练编码器中提取视觉概念,利用多模态大语言模型(MLLMs)根据视觉可辨识性和任务相关性对概念进行标注与过滤,并通过重建引导优化选择独立子集。借助CLIP的视觉-文本对齐能力,将图像表示分解为概念嵌入的线性组合,以适应CBMs框架。在11个图像分类任务上的大量实验表明,PCBM-ReD实现了最先进的准确率,缩小了与端到端模型的性能差距,并展现出更优的可解释性。

原文摘要 · Abstract (English)

Deep learning has achieved remarkable success in image recognition, yet their inherent opacity poses challenges for deployment in critical domains. Concept-based interpretations aim to address this by explaining model reasoning through human-understandable concepts. However, existing post-hoc methods and ante-hoc concept bottleneck models (CBMs), suffer from limitations such as unreliable concept relevance, non-visual or labor-intensive concept definitions, and model or data-agnostic assumptions. This paper introduces Post-hoc Concept Bottleneck Model via Representation Decomposition (PCBM-ReD), a novel pipeline that retrofits interpretability onto pretrained opaque models. PCBM-ReD automatically extracts visual concepts from a pre-trained encoder, employs multimodal large language models (MLLMs) to label and filter concepts based on visual identifiability and task relevance, and selects an independent subset via reconstruction-guided optimization. Leveraging CLIP's visual-text alignment, it decomposes image representations into linear combination of concept embeddings to fit into the CBMs abstraction. Extensive experiments across 11 image classification tasks show PCBM-ReD achieves state-of-the-art accuracy, narrows the performance gap with end-to-end models, and exhibits better interpretability.

可解释AI概念瓶颈视觉分解CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。