提出概念完整性新标准,提升模型可解释性与鲁棒性
Concept Labels Are Not Enough: Rethinking Concept Bottleneck Models through Representation Integrity
- 用组内一致性和覆盖度衡量概念表征的完整性
- 仅用10%标签仍保持93%性能,显著降低对标注依赖
- 适合关注可解释性、抗干扰能力的研究者
尽管深度神经网络预测性能强,其内部推理过程仍难以解析和控制。概念瓶颈模型(CBMs)通过人类可理解的概念分解预测,实现概念级可解释性与干预。然而,现有方法易受概念漂移和信息泄露影响,且缺乏对支撑概念的内部特征组织方式的评估,也未揭示失败背后的表征缺陷。本文提出‘概念完整性’这一关键属性:概念支持应形成语义连贯、不碎片化的功能组。为此引入组一致性(GC)与概念覆盖率(CC),合成概念完整性评分(CIS)。进一步提出概念完整性正则化(CIR),促使概念支撑滤波器形成清晰分离的组。所提隐式解耦概念瓶颈模型(LDCBM)无需区域标注即可应用CIR。在三个数据集上,CIS排名与概念及任务准确率排名不同,验证了其作为补充表征准则的有效性。仅使用10%概念标签,LDCBM即保持约93%全标签性能。背景掩码与概念干预实验表明,更高完整性对应更低对无关上下文敏感性与更强概念修正响应。
原文摘要 · Abstract (English)
Although deep neural networks achieve strong predictive performance, their internal reasoning often remains difficult to inspect and control. Concept Bottleneck Models (CBMs) address this opacity by factoring predictions through human-understandable concepts, thereby enabling concept-level inspection and intervention. However, CBMs remain vulnerable to concept shift and information leakage, while existing evaluations neither reveal how the internal features supporting each concept are organized nor identify the representation-level deficiency associated with these failures. We argue that this missing property is concept integrity: concept support should form a semantically coherent and non-fragmented functional group. To characterize this property, we propose group coherence (GC) and concept coverage (CC) as integrity components and aggregate them into the concept integrity score (CIS). We further introduce Concept Integrity Regularization (CIR) to encourage coherent and separated groups of concept-supporting filters. Our latent decoupled concept bottleneck model (LDCBM) applies CIR without requiring region annotations. Across three datasets, CIS rankings differ from rankings by concept and task accuracy, supporting concept integrity as a complementary representation-level criterion. Moreover, with only 10\% of concept labels, LDCBM retains approximately 93\% of its full-label task performance. Background-masking and concept-intervention further associate stronger integrity profiles with lower sensitivity to nuisance context and more responsive concept correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。