arXiv:2602.14626cs.LG2026-02被引 2

用信息瓶颈提升可解释模型的准确率与可信度

Concepts' Information Bottleneck Models

  • 在概念层引入信息瓶颈正则,控制输入与概念的冗余关联
  • 六类模型三数据集测试中,准确率和概念干预可靠性均提升
  • 无需改架构或额外标注,适配性强,适合追求可解释性的研究者

概念瓶颈模型(CBM)通过人类可理解的概念层实现可解释决策,但常面临准确率下降和概念泄露问题。本文提出一种显式的信息瓶颈正则,对输入与概念之间的互信息 $I(X;C)$ 进行惩罚,同时保留概念与任务输出间的相关性 $I(C;Y)$,促使概念表示达到最小充分性。推导出两种实用变体(变分目标与基于熵的代理),可无缝集成至标准CBM训练流程,无需修改架构或引入额外监督。在六类CBM模型与三个基准上的实验表明,加入信息瓶颈正则的模型始终优于原始版本。信息平面分析进一步验证了预期行为。结果表明,强制最小充分概念瓶颈能同时提升预测性能与概念级干预的可靠性。该正则提供了一种理论坚实、架构无关的路径,使CBM更可信且可干预,通过统一训练协议解决了以往评估不一致问题,并在多种模型与数据集上展现出稳健增益。

原文摘要 · Abstract (English)

Concept Bottleneck Models (CBMs) aim to deliver interpretable predictions by routing decisions through a human-understandable concept layer, yet they often suffer reduced accuracy and concept leakage that undermines faithfulness. We introduce an explicit Information Bottleneck regularizer on the concept layer that penalizes $I(X;C)$ while preserving task-relevant information in $I(C;Y)$, encouraging minimal-sufficient concept representations. We derive two practical variants (a variational objective and an entropy-based surrogate) and integrate them into standard CBM training without architectural changes or additional supervision. Evaluated across six CBM families and three benchmarks, the IB-regularized models consistently outperform their vanilla counterparts. Information-plane analyses further corroborate the intended behavior. These results indicate that enforcing a minimal-sufficient concept bottleneck improves both predictive performance and the reliability of concept-level interventions. The proposed regularizer offers a theoretic-grounded, architecture-agnostic path to more faithful and intervenable CBMs, resolving prior evaluation inconsistencies by aligning training protocols and demonstrating robust gains across model families and datasets.

可解释性信息瓶颈概念模型模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。