发现探针准确率不能可靠衡量概念对齐,提出新方法与评估指标。
Probing the Probes: Methods and Metrics for Concept Alignment
- 用空间线性归因法定位概念,改进传统探针的偏差问题。
- 误对齐探针也能达到高准确率,说明原指标不可靠。
- 提出三类量化对齐度的指标,适合模型可解释性研究者使用。
在可解释人工智能中,概念激活向量(CAVs)通常通过训练线性分类器探针,在深度神经网络的激活空间中检测人类可理解的概念。普遍认为探针准确率高即表示其忠实代表目标概念。然而我们发现,探针分类准确率本身并不能可靠衡量概念对齐程度——即向量是否真实捕捉了目标概念。事实上,探针更可能捕获的是虚假相关关系而非目标概念。我们的分析显示,刻意设计的误对齐探针(利用虚假相关)能达到与标准探针相近的准确率。为解决此问题,我们提出一种基于空间线性归因的概念定位新方法,并系统比较了其与现有特征可视化技术在检测和缓解概念错位方面的表现。进一步提出了三类量化概念对齐的指标:硬准确率、分割得分和增强鲁棒性。分析表明,具备平移不变性和空间对齐性的探针能显著提升概念对齐度。这些发现强调应采用基于对齐的评估指标,而非仅依赖探针准确率,并需根据模型架构和目标概念特性定制探针。
原文摘要 · Abstract (English)
In explainable AI, Concept Activation Vectors (CAVs) are typically obtained by training linear classifier probes to detect human-understandable concepts as directions in the activation space of deep neural networks. It is widely assumed that a high probe accuracy indicates a CAV faithfully representing its target concept. However, we show that the probe's classification accuracy alone is an unreliable measure of concept alignment, i.e., the degree to which a CAV captures the intended concept. In fact, we argue that probes are more likely to capture spurious correlations than they are to represent only the intended concept. As part of our analysis, we demonstrate that deliberately misaligned probes constructed to exploit spurious correlations, achieve an accuracy close to that of standard probes. To address this severe problem, we introduce a novel concept localization method based on spatial linear attribution, and provide a comprehensive comparison of it to existing feature visualization techniques for detecting and mitigating concept misalignment. We further propose three classes of metrics for quantitatively assessing concept alignment: hard accuracy, segmentation scores, and augmentation robustness. Our analysis shows that probes with translation invariance and spatial alignment consistently increase concept alignment. These findings highlight the need for alignment-based evaluation metrics rather than probe accuracy, and the importance of tailoring probes to both the model architecture and the nature of the target concept.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。