arXiv:2601.21944cs.LG2026-01

提出可衡量解释性与灵活性权衡的新指标Clarity,揭示稀疏模型的决策机制。

Clarity: The Flexibility-Interpretability Trade-Off in Sparsity-aware Concept Bottleneck Models

  • 引入Clarity指标,量化概念激活的稀疏性与精确性对解释性的影响。
  • 发现不同方法在相似性能下解释性差异显著,存在灵活与可解释的权衡。
  • 人类实验验证Clarity比传统指标更贴近真实信任度,适合评估可信AI。

深度学习在计算机视觉中的广泛应用加剧了对模型可解释性的关注。尽管性能优异,这些模型常被视为黑箱,其决策过程缺乏系统性研究。现有解释性方法虽多,但对依赖稀疏性来“诱导”可解释性的方法,仍缺乏客观评估。本文研究概念瓶颈模型(CBMs)中建模选择如何影响概念表示的语义对齐。提出一种后验诊断指标Clarity,捕捉下游性能与概念激活稀疏性、精确性之间的相互作用。基于带有真实概念标注的数据集,评估了基于视觉语言模型(VLM)和属性预测器的CBMs,在三种稀疏化策略(ℓ₁、ℓ₀、伯努利)及多种文献中常用稀疏感知CBM方法下的表现。实验揭示关键的灵活性-解释性权衡:模型为优化任务性能而偏离语义对齐的能力。结果显示,即使在相近性能水平下,不同方法行为差异明显。最终通过严谨的人类研究验证,Clarity与人类信任度相关性显著高于标准评估指标。

原文摘要 · Abstract (English)

The widespread adoption of deep learning models in computer vision has intensified concerns about interpretability. Despite strong performance, these models are often treated as black boxes, with limited systematic investigation of their decision-making processes. While many interpretability methods exist, objective evaluation of learned representations remains limited, particularly for approaches that rely on sparsity to "induce" interpretability. In this work, we investigate how modeling choices in Concept Bottleneck Models (CBMs) affect the semantic alignment of concept representations. We introduce Clarity, a post-hoc diagnostic measure that captures the interplay between downstream performance and the sparsity and precision of concept activations. Using an interpretability assessment framework grounded in datasets with ground-truth concept annotations, we evaluate both VLM- and attribute predictor-based CBMs across three amortized sparsity-inducing strategies ($\ell_1$, $\ell_0$, and Bernoulli-based), alongside several widely used sparsity-aware CBM methods from the literature. Our experiments reveal a critical flexibility-interpretability trade-off: a model's capacity to optimize task performance by deviating from semantic alignment. We demonstrate that under this trade-off, different methods exhibit markedly different behaviors even at comparable performance levels. Finally, we validate our framework through a principled human study, which confirms that Clarity aligns significantly more closely with human trust than standard evaluation metrics.

可解释性稀疏性概念瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。