用多模型委员会量化评估概念瓶颈模型解释质量。
CBX-Bench: A Human-Aligned MLLM Council for Benchmarking Concept Bottleneck Model Explanations

- 构建由五个开源多模态大模型组成的委员会,自动评分图像解释质量。
- 在900个图像对比任务中,委员会对一致人类判断的准确率达83%。
- 适合关注可解释性评估的模型开发者与研究者使用。
概念瓶颈模型(CBMs)旨在通过人类可理解的概念表达视觉分类结果,以提升可解释性。然而当前评估仍主要依赖下游分类准确率,辅以零散的定性示例,缺乏量化指标。由于大规模真实概念标注不可行,且概念列表尚无共识,这一问题尤为突出。为此,我们构建了一个多模态大模型(MLLM)委员会:给定一张图像及其CBM解释,委员会生成解释质量评分。为验证其有效性,我们开展人类研究,在CUB-200、ImageNet-100和Places365上对900个图像对比项进行2700次判断,确定人类参考标准。基于此,我们的五模型委员会在严格人类偏好排序中恢复超过70%,在人类一致判断项中达83%。在此基础上,我们推出公开基准测试平台CBX-Bench,支持新CBM提交解释并获得委员会评分与数据集级排名,实现超越准确率与孤立示例的人类对齐、可扩展的解释质量评估。项目地址:https://github.com/meric-karadag/cbx-bench。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human-understandable concepts. Although interpretability is the central motivation for CBMs, they are still largely evaluated as predictive models by downstream classification accuracy, supplemented by isolated qualitative examples. This highlights a pressing need for quantitative measures, a challenge complicated by the infeasibility of ground-truth concept annotation at scale and the open nature of concept lists due to a lack of consensus. To fill this gap, we develop a multimodal large language model (MLLM) council that, given an image and its CBM explanation, produces an explanation quality score. To ground and validate the council, we first conduct a human study to establish a ground-truth reference for CBM explanation quality: for an image, annotators compare explanations from two of LF-CBM, VLG-CBM, and CBM-Suite and choose the more useful one, or mark them as equally good or equally bad, yielding 2700 judgments over 900 image-comparison items on CUB-200, ImageNet-100, and Places365. Against this human reference, our five-model council, consisting of open-weight MLLMs, recovers over 70% of strict human preference rankings, rising to 83% on items where human annotators unanimously agree. Building on this validated council, we introduce CBX-Bench, a public benchmark and leaderboard: authors of new CBMs can submit their model's explanations, and CBX-Bench scores them with the council and maintains dataset-level rankings of explanation quality. CBX-Bench thus provides a human-aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples. The benchmark is available at https://github.com/meric-karadag/cbx-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。