用语言模型让概念瓶颈模型支持自然语言干预,提升可解释性与交互性。
Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
- 用语言模型替代传统数值分类器,基于概念语义推理
- 在9个数据集上表现更优,且支持概念增删与知识注入
- 适合需要高可解释性与灵活干预的场景
概念瓶颈模型(CBMs)通过先预测人类可理解的概念,再通过简单分类器映射到标签,提供内在可解释性。传统CBM使用固定线性分类器对概念得分进行处理,干预方式仅限于手动调整数值,无法在测试时添加新概念或融入领域知识,尤其在无监督设置下因概念激活噪声大、密集而难以有效干预。我们提出Chat-CBM,将基于分数的分类器替换为直接对概念语义进行推理的语言模型分类器。通过将预测锚定在概念的语义空间,Chat-CBM在保持CBM可解释性的前提下,支持概念修正、新增/删除概念、引入外部知识及高层推理引导等更丰富直观的干预方式。借助冻结的大语言模型的语言理解与少样本能力,Chat-CBM扩展了CBM的干预接口,超越数值编辑,并在无监督场景中仍保持有效性。九个数据集上的实验表明,Chat-CBM不仅预测性能更高,还显著提升了用户交互体验。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) provide inherent interpretability by first predicting a set of human-understandable concepts and then mapping them to labels through a simple classifier. While users can intervene in the concept space to improve predictions, traditional CBMs typically employ a fixed linear classifier over concept scores, which restricts interventions to manual value adjustments and prevents the incorporation of new concepts or domain knowledge at test time. These limitations are particularly severe in unsupervised CBMs, where concept activations are often noisy and densely activated, making user interventions ineffective. We introduce Chat-CBM, which replaces score-based classifiers with a language-based classifier that reasons directly over concept semantics. By grounding prediction in the semantic space of concepts, Chat-CBM preserves the interpretability of CBMs while enabling richer and more intuitive interventions, such as concept correction, addition or removal of concepts, incorporation of external knowledge, and high-level reasoning guidance. Leveraging the language understanding and few-shot capabilities of frozen large language models, Chat-CBM extends the intervention interface of CBMs beyond numerical editing and remains effective even in unsupervised settings. Experiments on nine datasets demonstrate that Chat-CBM achieves higher predictive performance and substantially improves user interactivity while maintaining the concept-based interpretability of CBMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。