arXiv:2506.13831cs.LGcs.AI2025-06被引 6

为CLIP嵌入中的概念结构提供统计验证,提升可解释性与鲁棒性。

Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation

  • 基于假设检验识别嵌入空间中旋转敏感的结构
  • 重构误差更低,概念更稳定且可复现
  • 适合需要可靠解释性的图像模型研究者

概念化方法旨在从深度神经网络(如CLIP)的内部表示中识别人类可理解的概念,是解释模型行为的有力手段。然而现有方法缺乏统计严谨性,难以验证所发现概念的有效性或比较不同方法。为此,我们提出一种假设检验框架,用于量化CLIP嵌入空间中的旋转敏感结构。一旦识别出此类结构,我们进一步设计后验概念分解方法,该方法具有理论保证,确保发现的概念是稳健、可复现的模式而非方法特异性伪影,并在重构误差上优于现有技术。实证表明,该算法在重建精度与概念可解释性之间取得良好平衡,有助于消除数据中的虚假线索。应用于一个典型的虚假相关性数据集后,去除虚假背景概念使最差组准确率提升了22.6%。

原文摘要 · Abstract (English)

Concept-based approaches, which aim to identify human-understandable concepts within a model's internal representations, are a promising method for interpreting embeddings from deep neural network models, such as CLIP. While these approaches help explain model behavior, current methods lack statistical rigor, making it challenging to validate identified concepts and compare different techniques. To address this challenge, we introduce a hypothesis testing framework that quantifies rotation-sensitive structures within the CLIP embedding space. Once such structures are identified, we propose a post-hoc concept decomposition method. Unlike existing approaches, it offers theoretical guarantees that discovered concepts represent robust, reproducible patterns (rather than method-specific artifacts) and outperforms other techniques in terms of reconstruction error. Empirically, we demonstrate that our concept-based decomposition algorithm effectively balances reconstruction accuracy with concept interpretability and helps mitigate spurious cues in data. Applied to a popular spurious correlation dataset, our method yields a 22.6% increase in worst-group accuracy after removing spurious background concepts.

可解释性概念分解CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。