arXiv:2504.14094cs.LGcs.AI2025-04被引 14

提出量化概念模型泄漏的新框架,提升高风险场景下的可解释性。

Leakage and Interpretability in Concept-Based Models

  • 用信息论定义两种泄漏指标:概念-任务与概念间泄漏。
  • 实验证明泄漏指标能精准预测模型干预后的行为变化。
  • 揭示概念嵌入模型中的多重泄漏问题,提供降漏指南。

概念基础模型通过预测高层中间概念来提升可解释性,是高风险场景下有前景的部署方法。然而,这类模型易出现信息泄漏,即利用学习概念中编码的非预期信息。本文提出一种信息论框架,严格刻画并量化泄漏,定义两个互补度量:概念-任务泄漏(CTL)和概念间泄漏(ICL)得分。实验表明,这些度量对模型在干预下的行为具有强预测能力,优于现有方法。基于该框架,我们识别出泄漏的主要成因,并以概念嵌入模型为例,发现除设计固有的概念-任务泄漏外,还存在概念间泄漏和对齐泄漏。最后,提出一套实用设计准则,用于降低泄漏、保障可解释性。

原文摘要 · Abstract (English)

Concept-based Models aim to improve interpretability by predicting high-level intermediate concepts, representing a promising approach for deployment in high-risk scenarios. However, they are known to suffer from information leakage, whereby models exploit unintended information encoded within the learned concepts. We introduce an information-theoretic framework to rigorously characterise and quantify leakage, and define two complementary measures: the concepts-task leakage (CTL) and interconcept leakage (ICL) scores. We show that these measures are strongly predictive of model behaviour under interventions and outperform existing alternatives. Using this framework, we identify the primary causes of leakage and, as a case study, analyse how it manifests in Concept Embedding Models, revealing interconcept and alignment leakage in addition to the concepts-task leakage present by design. Finally, we present a set of practical guidelines for designing concept-based models to reduce leakage and ensure interpretability.

可解释性信息泄漏概念模型信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。