构建可控制变量的合成数据集,评估概念瓶颈模型性能。
Measuring What Matters: Synthetic Benchmarks for Concept Bottleneck Models

- 设计可控合成基准,模拟不同数据模态与概念选择。
- 验证多种概念瓶颈模型在不同标注质量下的表现差异。
- 帮助诊断模型失效原因,适合研究可解释性算法者使用。
概念瓶颈模型通过输入中检测到的高层概念预测结果。尽管概念能带来可解释性优势,但多数数据集缺乏概念标签,限制了研究人员判断哪些问题适合此类模型、分离影响其性能的因素或发现失败原因的能力。本文提出针对概念瓶颈模型的合成基准,聚焦两大应用场景:决策支持(辅助人类决策)和自动化(无需监督处理常规任务)。这些基准可在控制数据模态、概念选择、标注质量和完整性等关键因素的前提下生成带标签数据集。我们展示了如何用这些基准评估代表性概念瓶颈模型,并证明其可有效诊断模型失败模式,指导后续测试。
原文摘要 · Abstract (English)
Concept bottleneck models predict outcomes from high-level concepts detected in inputs. Although concepts provide a simple way to reap benefits from interpretability, very few datasets include concept labels. This limits researchers' ability to determine which problems are suitable for these models, isolate the factors that drive their performance or lead to failures, or uncover which algorithms perform well. In this paper, we develop synthetic benchmarks for concept-bottleneck models, focusing on their two main use cases: decision support, in which models assist humans in making better decisions, and automation, in which models handle routine tasks without supervision. Our benchmarks can generate labeled datasets while controlling for properties that affect performance, including data modality, concept choice, annotation quality, and completeness. We demonstrate how the benchmarks can be used to evaluate representative classes of concept bottleneck models. Our demonstrations show how the benchmarks can diagnose failure modes and guide follow-up testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。