提出量化模型可解释性与准确率权衡的新方法,让模型既准又易懂。
Quantifying the Accuracy-Interpretability Trade-Off in Concept-Based Sidechannel Models
- 构建统一的概率侧通道元模型,涵盖现有方法
- 引入侧通道独立性评分,衡量模型对不可解释信息的依赖程度
- 通过正则化提升可解释性,适合需要透明决策的场景
概念瓶颈模型(CBNMs)通过强制预测仅基于人类可理解的概念来实现可解释性,但限制了信息流,常导致准确率下降。概念侧通道模型(CSMs)通过引入绕过瓶颈的侧通道传输任务相关额外信息,提升了准确率,但牺牲了可解释性,因为预测可能依赖于不可解释的表示。目前尚无系统方法控制这一根本权衡。本文首次填补此空白:首先提出统一的概率概念侧通道元模型,涵盖现有模型;在此基础上,提出侧通道独立性评分(SIS),通过对比有无侧通道信息时的预测差异,量化模型对侧通道的依赖;进一步提出SIS正则化,显式惩罚对侧通道的依赖以增强可解释性;最后分析预测器表达能力与侧通道依赖共同决定可解释性的机制,揭示不同架构下的内在权衡。实验证明,仅追求准确率的先进CSMs表现出低可解释性表示,而SIS正则化显著提升其可解释性、可控性和可解释预测器的质量。本工作为开发兼顾准确率与可解释性的模型提供了理论与实践工具。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBNMs) are deep learning models that provide interpretability by enforcing a bottleneck layer where predictions are based exclusively on human-understandable concepts. However, this constraint also restricts information flow and often results in reduced predictive accuracy. Concept Sidechannel Models (CSMs) address this limitation by introducing a sidechannel that bypasses the bottleneck and carry additional task-relevant information. While this improves accuracy, it simultaneously compromises interpretability, as predictions may rely on uninterpretable representations transmitted through sidechannels. Currently, there exists no principled technique to control this fundamental trade-off. In this paper, we close this gap. First, we present a unified probabilistic concept sidechannel meta-model that subsumes existing CSMs as special cases. Building on this framework, we introduce the Sidechannel Independence Score (SIS), a metric that quantifies a CSM's reliance on its sidechannel by contrasting predictions made with and without sidechannel information. We propose SIS regularization, which explicitly penalizes sidechannel reliance to improve interpretability. Finally, we analyze how the expressivity of the predictor and the reliance of the sidechannel jointly shape interpretability, revealing inherent trade-offs across different CSM architectures. Empirical results show that state-of-the-art CSMs, when trained solely for accuracy, exhibit low representation interpretability, and that SIS regularization substantially improves their interpretability, intervenability, and the quality of learned interpretable task predictors. Our work provides both theoretical and practical tools for developing CSMs that balance accuracy and interpretability in a principled manner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。