arXiv:2409.17568cs.AI2024-09

研究多标签模型中让分类器输出尽可能多标签的对抗攻击方法。

Showing Many Labels in Multi-label Classification Models: An Empirical Study of Adversarial Examples

  • 提出'显示更多标签'攻击,目标是最大化预测标签数量。
  • 迭代攻击在8种场景下成功率显著高于单步攻击,可触发全部标签。
  • 适用于评估多标签模型安全性的研究人员与安全防御开发者。

随着深度神经网络的快速发展,其已广泛应用于多个领域。然而研究表明,深度神经网络易受对抗样本影响,这一现象在多标签任务中同样存在。为深入探究多标签对抗样本,本文引入一种新型攻击方式——'显示更多标签',旨在最大化分类器预测结果中的标签数量。我们选取九种攻击算法,在四种主流多标签数据集(VOC2007、VOC2012、NUS-WIDE、COCO)上,对两种目标模型(ML-LIW 和 ML-GCN)进行测试。在八个不同场景中记录各算法成功展示预期标签数量的比率。实验表明,在'显示更多标签'攻击下,迭代攻击性能显著优于单步攻击,且在某些情况下可使模型输出全部标签。

原文摘要 · Abstract (English)

With the rapid development of Deep Neural Networks (DNNs), they have been applied in numerous fields. However, research indicates that DNNs are susceptible to adversarial examples, and this is equally true in the multi-label domain. To further investigate multi-label adversarial examples, we introduce a novel type of attacks, termed "Showing Many Labels". The objective of this attack is to maximize the number of labels included in the classifier's prediction results. In our experiments, we select nine attack algorithms and evaluate their performance under "Showing Many Labels". Eight of the attack algorithms were adapted from the multi-class environment to the multi-label environment, while the remaining one was specifically designed for the multi-label environment. We choose ML-LIW and ML-GCN as target models and train them on four popular multi-label datasets: VOC2007, VOC2012, NUS-WIDE, and COCO. We record the success rate of each algorithm when it shows the expected number of labels in eight different scenarios. Experimental results indicate that under the "Showing Many Labels", iterative attacks perform significantly better than one-step attacks. Moreover, it is possible to show all labels in the dataset.

多标签对抗攻击安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。