提出新模型让概念瓶颈更真实,提升可解释性与干预可靠性。
There Was Never a Bottleneck in Concept Bottleneck Models
- 引入信息瓶颈正则化,强制每个组件只保留相关概念信息。
- 实验证明新模型在概念解耦上优于传统CBM,干预效果更可信。
- 适合关注模型可解释性与可控生成的研究者使用。
深度学习表示通常难以解释,限制了其在敏感场景的应用。概念瓶颈模型(CBMs)通过学习支持目标任务且每个组件预测预定义概念的表示来缓解此问题。本文指出,传统CBM并不存在真正的瓶颈:组件能预测概念,并不意味着它仅编码该概念信息。这一缺陷影响可解释性及干预有效性。为此,我们提出最小概念瓶颈模型(MCBMs),通过在训练损失中加入信息瓶颈(IB)正则化项,约束每个表示组件仅保留与其对应概念相关的信息。该方法基于变分正则化实现,使MCBMs获得更具可解释性的表示,支持合理的概念级干预,并与概率理论基础一致。
原文摘要 · Abstract (English)
Deep learning representations are often difficult to interpret, which can hinder their deployment in sensitive applications. Concept Bottleneck Models (CBMs) have emerged as a promising approach to mitigate this issue by learning representations that support target task performance while ensuring that each component predicts a concrete concept from a predefined set. In this work, we argue that CBMs do not impose a true bottleneck: the fact that a component can predict a concept does not guarantee that it encodes only information about that concept. This shortcoming raises concerns regarding interpretability and the validity of intervention procedures. To overcome this limitation, we propose Minimal Concept Bottleneck Models (MCBMs), which incorporate an Information Bottleneck (IB) objective to constrain each representation component to retain only the information relevant to its corresponding concept. This IB is implemented via a variational regularization term added to the training loss. As a result, MCBMs yield more interpretable representations, support principled concept-level interventions, and remain consistent with probability-theoretic foundations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。