用语义概念定位模型错误群体,解释更精准。
CB-SLICE: Concept-Based Interpretable Error Slice Discovery

- 基于概念瓶颈模型,从语义概念误判定位错误样本
- 在多个数据集上发现典型偏差,解释比现有方法更准确
- 适合调试模型、分析偏见,尤其关注可解释性的人
尽管深度学习模型在平均性能上表现良好,但在特定人群群体中常出现系统性错误,称为错误切片(error slices)。识别这些群体及其失败根源对模型调试和偏见缓解至关重要。然而,现有错误切片发现方法(SDMs)生成的解释与模型推理过程脱节,仅近似底层错误原因,可能不准确。本文提出CB-SLICE,利用概念瓶颈模型(CBMs),其预测直接依赖于人类可理解的语义概念。由于下游任务失败常源于概念误判,概念表示成为错误切片识别的理想候选,能提供与错误源直接关联的细粒度解释。基于此,我们构建了基于概念的SDM,将具有共同概念预测失败的样本分组,并识别每个切片失败模式中最关键的概念。在多个基准测试中,CB-SLICE优于当前最佳方法,不仅能更有效发现已知偏见,还提供了更丰富、更可信的模型错误解释。
原文摘要 · Abstract (English)
Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Identifying these groups and the root causes of their failures is critical for model debugging and bias mitigation. However, existing error Slice Discovery Methods (SDMs) typically generate explanations disconnected from the model's inference process, thus only approximating the underlying error source and may be inaccurate. We address this limitation by leveraging Concept Bottleneck Models (CBMs), whose predictions are directly dependent on human-understandable semantic concepts. Since downstream task failures in CBMs commonly arise from concept mispredictions, concept representations provide a strong candidate for error slice identification, offering fine-grained explanations directly linked to the error source. Building on this insight, we introduce CB-SLICE, a concept-based SDM that groups samples with shared concept prediction failures and identifies the keyword concepts most responsible for each slice's failure mode. Across multiple benchmarks, we show that CB-SLICE outperforms state-of-the-art methods in uncovering well-known biases while providing richer and more faithful explanations of model errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。