提出评估概念检测可靠性的方法,解决解释不可靠问题。
Assessing Reliability of Symbol Detection in Concept Bottleneck Models

- 通过独立训练检测器与分类头互换,检验符号可靠性
- 监督减弱时任务准确率不变,但交换性能降至随机水平
- 引入可靠性感知训练,显著提升不可靠情况下的鲁棒性
概念瓶颈模型(CBMs)是可解释AI的重要工具,其预测依赖于人类可理解的符号。然而高任务准确率不等于符号检测可靠:联合训练的CBMs可能在瓶颈层编码任务特异性捷径,导致解释不可靠。本文通过交换独立训练的概念检测器与分类头(共享符号词汇),利用性能下降、概念级指标和符号不确定性估计,识别易产生虚假激活的概念。进一步提出一种可靠性感知训练策略:共享检测器同时优化多个分类头,并对依赖全局或实例级不可靠符号施加惩罚。在全监督的CUB-200-2011数据集上,检测器与分类头几乎可自由互换(互换损失低于1个百分点,相对保留率高于99%,无概念低于随机水平);而在受控合成任务中,随着概念监督权重降低,模型维持近完美任务准确率,但互换准确率与真实概念一致性骤降至随机水平。所提方法显著缓解该泄漏现象,在易泄露场景下使互换准确率大致翻倍。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) are a relevant tool for explainable Artificial Intelligence because they make their predictions through human-interpretable symbols. However, high task accuracy does not guarantee that these symbols are detected faithfully: jointly trained CBMs may encode task-specific shortcuts in the bottleneck, making their explanations unreliable. In this paper, we study concept-detection reliability by swapping independently trained concept detectors and classification heads that share the same symbolic vocabulary. We use the resulting performance degradation, concept-level metrics, and symbol-wise uncertainty estimates to identify concepts that are especially prone to spurious firing. Finally, we propose a reliability-aware training strategy in which a shared concept detector is optimized with multiple classification heads and penalized for relying on globally or instance-wise unreliable symbols. On CUB-200-2011 with full concept supervision, detectors and heads are almost freely interchangeable (swap drop below one accuracy point, relative retention above $99\%$, and no concept detected below chance), whereas on a controlled synthetic task we show that, as the concept-supervision weight is reduced, models keep near-perfect task accuracy while swapped accuracy and agreement with the ground-truth concepts collapse to chance. Our reliability-aware training substantially mitigates this leakage, roughly doubling swap accuracy in the leaky regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。