让概念模型可验证,看清每处判断的视觉依据
Towards Fine-Grained and Verifiable Concept Bottleneck Models

- 为每个概念绑定局部视觉证据,实现精准定位
- 在医学影像上达到与传统CBM相当的准确率
- 支持人工核查概念是否真实对应图像内容
概念瓶颈模型(CBMs)通过引入人类可理解的概念提升黑箱预测的可解释性。然而现有方法难以验证预测概念是否真正对应视觉证据,影响可靠性。本文提出细粒度CBM框架,将每个概念与局部视觉证据关联,使用户可直接检查概念编码的位置与方式。实验表明,该方法在医学影像基准上实现了信息完备的概念空间,预测性能媲美标准CBM,同时显著提升透明度。不同于事后归因方法,本框架可验证概念表征的存在性与正确性,弥合可解释性与可验证性之间的鸿沟。该方法增强了CBM的信任度,建立了人-模型在概念层面的可靠交互机制,推动更可信、临床可用的概念驱动学习系统发展。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) offer interpretable alternatives to black-box predictors by introducing human-relatable concepts before the final output. However, existing CBMs struggle to verify whether predicted concepts correspond to the correct visual evidence, limiting their reliability. We propose a fine-grained CBM framework that grounds each concept in localized visual evidence, enabling direct inspection of where and how concepts are encoded. This design allows users to interpret predictions and verify that the model learns intended concepts rather than spurious correlations. Experiments on medical imaging benchmarks show that our learned concept space is information-complete and achieves predictive performance comparable to standard CBMs, while substantially improving transparency. Unlike post-hoc attribution methods, our framework validates both the presence and correctness of concept representations, bridging interpretability with verifiability. Our approach enhances the trustworthiness of CBMs and establishes a principled mechanism for human-model interaction at the concept level, paving the way toward more reliable and clinically actionable concept-based learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。