用人类能理解的概念检测模型失败,更准还更透明。
Interpretable Failure Detection with Human-Level Concepts
- 基于概念激活排序判断模型是否出错,比传统方法更细致。
- 在ImageNet上误报率降3.7%,EuroSAT上降9%。
- 适合需要高可靠性和可解释性的安全关键场景。
可靠故障检测在安全关键应用中至关重要。然而,神经网络对分类错误的样本常产生过度自信的预测。现有置信度评分方法依赖类别级信号(如logits)进行故障检测,存在缺陷。本文提出一种新策略,利用人类可理解的概念实现双重目标:准确检测模型失败,并透明解释原因。通过整合每个类别的多维信号,方法可实现更细粒度的置信度评估。我们提出一种基于概念激活排序的简单但高效的方法,无需复杂组件,在多个真实世界图像分类基准上显著降低误报率:ImageNet上减少3.7%,EuroSAT上减少9%。
原文摘要 · Abstract (English)
Reliable failure detection holds paramount importance in safety-critical applications. Yet, neural networks are known to produce overconfident predictions for misclassified samples. As a result, it remains a problematic matter as existing confidence score functions rely on category-level signals, the logits, to detect failures. This research introduces an innovative strategy, leveraging human-level concepts for a dual purpose: to reliably detect when a model fails and to transparently interpret why. By integrating a nuanced array of signals for each category, our method enables a finer-grained assessment of the model's confidence. We present a simple yet highly effective approach based on the ordinal ranking of concept activation to the input image. Without bells and whistles, our method significantly reduce the false positive rate across diverse real-world image classification benchmarks, specifically by 3.7% on ImageNet and 9% on EuroSAT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。