arXiv:2411.19853cs.LGcs.CV2024-11被引 3

分析分类模型在对抗攻击下的类别差异脆弱性,发现误报数决定类别的易攻破程度。

Towards Class-wise Robustness Analysis

  • 按类别分析对抗训练模型的鲁棒性差异,关注误报率影响
  • 发现特定类别的误报数量显著提升其被攻击风险
  • 提出类别误报得分评估机制,适合安全敏感场景研究者

尽管深度神经网络在众多下游任务中表现优异,但其在真实场景中的应用受限于对领域偏移(如常见数据损坏和对抗攻击)的敏感性。对抗样本与数据损坏会显著降低深度分类模型的性能。研究人员已开发出多种鲁棒神经架构以增强分类器决策能力,但多数工作依赖有效的对抗训练方法,且主要关注整体模型鲁棒性,忽视了类别间鲁棒性的差异,而这一差异至关重要。利用弱鲁棒类别是攻击者误导图像识别模型的潜在途径。因此,本研究探讨对抗训练鲁棒分类模型中类别间的偏差,理解其隐空间结构,并分析各类别的强弱鲁棒特性。我们进一步评估了类别在常见数据损坏和对抗攻击下的鲁棒性,认识到类别脆弱性不仅体现在正确分类数量上。通过分析类别误报得分(Class False Positive Score),我们提出了一个公平评估每个类别被误分类敏感度的方法。

原文摘要 · Abstract (English)

While being very successful in solving many downstream tasks, the application of deep neural networks is limited in real-life scenarios because of their susceptibility to domain shifts such as common corruptions, and adversarial attacks. The existence of adversarial examples and data corruption significantly reduces the performance of deep classification models. Researchers have made strides in developing robust neural architectures to bolster decisions of deep classifiers. However, most of these works rely on effective adversarial training methods, and predominantly focus on overall model robustness, disregarding class-wise differences in robustness, which are critical. Exploiting weakly robust classes is a potential avenue for attackers to fool the image recognition models. Therefore, this study investigates class-to-class biases across adversarially trained robust classification models to understand their latent space structures and analyze their strong and weak class-wise properties. We further assess the robustness of classes against common corruptions and adversarial attacks, recognizing that class vulnerability extends beyond the number of correct classifications for a specific class. We find that the number of false positives of classes as specific target classes significantly impacts their vulnerability to attacks. Through our analysis on the Class False Positive Score, we assess a fair evaluation of how susceptible each class is to misclassification.

鲁棒性分析对抗攻击类别差异误报率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。