arXiv:2503.09068cs.LGcs.AI2025-03被引 1

无需标签信息,通过反事实样本揭示分类器漏洞。

Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information

  • 设计探针模型判断分类正确性,生成反事实样本。
  • 在MNIST上成功检测出分类器误判,准确率超90%。
  • 适合研究模型鲁棒性与无标签环境下的安全评估。

为提升深度神经网络分类器的可信度与透明性,需能解释其决策过程。现有实例级解释方法(如归因技术)在分析误分类时依赖人工干预,且跨类别分析易受主观偏见影响。本文提出一种新框架,通过反事实样本揭示分类器弱点。引入探针模型以二进制形式判断分类正确与否(命中/未命中),据此生成针对性反事实样本。在图像分类基准数据集上验证了探针对误分类的检测性能。进一步地,通过生成穿透探针的反事实样本,证明该框架可在不依赖标签信息的情况下,在MNIST数据集上有效识别目标分类器的脆弱点。

原文摘要 · Abstract (English)

To improve trust and transparency, it is crucial to be able to interpret the decisions of Deep Neural classifiers (DNNs). Instance-level examinations, such as attribution techniques, are commonly employed to interpret the model decisions. However, when interpreting misclassified decisions, human intervention may be required. Analyzing the attribu tions across each class within one instance can be particularly labor intensive and influenced by the bias of the human interpreter. In this paper, we present a novel framework to uncover the weakness of the classifier via counterfactual examples. A prober is introduced to learn the correctness of the classifier's decision in terms of binary code-hit or miss. It enables the creation of the counterfactual example concerning the prober's decision. We test the performance of our prober's misclassification detection and verify its effectiveness on the image classification benchmark datasets. Furthermore, by generating counterfactuals that penetrate the prober, we demonstrate that our framework effectively identifies vulnerabilities in the target classifier without relying on label information on the MNIST dataset.

模型解释反事实推理无标签检测鲁棒性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。