提出可解释性强的混合概念模型,提升准确率与简洁性。
Partially Shared Concept Bottleneck Models
- 用语言模型与视觉样本结合生成更精准的概念
- 通过激活模式合并概念,准确率提升1.0%-7.4%,概念减少
- 新指标联合评估准确率与概念简洁性,适合需要可解释性的场景
概念瓶颈模型(CBMs)通过在输入与预测间引入人类可理解的概念层来提升可解释性。尽管近期方法利用大语言模型(LLMs)和视觉语言模型(VLMs)自动生成概念,但仍面临三大挑战:视觉定位差、概念冗余、缺乏平衡预测准确率与概念紧凑性的原则性度量。本文提出部分共享概念瓶颈模型(PS-CBM),包含三个核心组件:(1) 多模态概念生成器,融合LLM语义与示例化视觉线索;(2) 部分共享概念策略,基于激活模式合并概念,兼顾特异性与紧凑性;(3) 概念高效准确率(CEA),一种后验度量,联合捕捉预测准确率与概念紧凑性。在十一组多样化数据集上的大量实验表明,PS-CBM持续优于当前最优的CBMs,分类准确率提升1.0%-7.4%,CEA提升2.0%-9.5%,且所需概念显著减少。结果验证了其在实现高准确率与强可解释性方面的有效性。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) enhance interpretability by introducing a layer of human-understandable concepts between inputs and predictions. While recent methods automate concept generation using Large Language Models (LLMs) and Vision-Language Models (VLMs), they still face three fundamental challenges: poor visual grounding, concept redundancy, and the absence of principled metrics to balance predictive accuracy and concept compactness. We introduce PS-CBM, a Partially Shared CBM framework that addresses these limitations through three core components: (1) a multimodal concept generator that integrates LLM-derived semantics with exemplar-based visual cues; (2) a Partially Shared Concept Strategy that merges concepts based on activation patterns to balance specificity and compactness; and (3) Concept-Efficient Accuracy (CEA), a post-hoc metric that jointly captures both predictive accuracy and concept compactness. Extensive experiments on eleven diverse datasets show that PS-CBM consistently outperforms state-of-the-art CBMs, improving classification accuracy by 1.0%-7.4% and CEA by 2.0%-9.5%, while requiring significantly fewer concepts. These results underscore PS-CBM's effectiveness in achieving both high accuracy and strong interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。