arXiv:2504.09459cs.LGcs.AI2025-04被引 12

提出量化概念模型信息泄露的理论方法,揭示维度与分类器对透明性的影响。

Measuring Leakage in Concept-Based Methods: An Information Theoretic Approach

  • 用信息论衡量概念嵌入中隐藏的非目标信息量
  • 发现特征维数和概念维数越高,泄露越严重
  • 推荐使用XGBoost评估,适合关注模型可解释性的研究者

概念瓶颈模型(CBMs)通过人类可理解的概念结构化预测以提升可解释性,但预测信号绕过概念瓶颈导致的信息泄露会损害其透明性。本文提出一种信息论方法,量化CBMs中概念嵌入所包含的非预期额外信息程度。通过受控的合成实验验证该方法的有效性,结果表明特征和概念维度显著影响泄露水平,且分类器选择影响测量稳定性,其中XGBoost表现最可靠。初步实验还显示该方法在软联合CBMs上呈现预期行为,表明其在真实场景中的潜在适用性。尽管当前研究聚焦于合成数据,未来可拓展至真实数据集。

原文摘要 · Abstract (English)

Concept Bottleneck Models (CBMs) aim to enhance interpretability by structuring predictions around human-understandable concepts. However, unintended information leakage, where predictive signals bypass the concept bottleneck, compromises their transparency. This paper introduces an information-theoretic measure to quantify leakage in CBMs, capturing the extent to which concept embeddings encode additional, unintended information beyond the specified concepts. We validate the measure through controlled synthetic experiments, demonstrating its effectiveness in detecting leakage trends across various configurations. Our findings highlight that feature and concept dimensionality significantly influence leakage, and that classifier choice impacts measurement stability, with XGBoost emerging as the most reliable estimator. Additionally, preliminary investigations indicate that the measure exhibits the anticipated behavior when applied to soft joint CBMs, suggesting its reliability in leakage quantification beyond fully synthetic settings. While this study rigorously evaluates the measure in controlled synthetic experiments, future work can extend its application to real-world datasets.

可解释性信息论概念模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。