arXiv:2502.19757cs.CVcs.CR2025-02

让交通标志识别模型误判,用明显干扰却人眼不受影响的攻击方式。

Snowball Adversarial Attack on Traffic Sign Classification

  • 利用人类视觉抗干扰优势,设计肉眼可见但误导模型的攻击。
  • 在多种图像上使顶尖交通标志识别模型性能大幅下降。
  • 揭示深度模型脆弱性,适合关注模型安全的研究者阅读。

机器学习模型的对抗攻击通常依赖微小、人眼难以察觉的扰动来误导分类器。这类策略旨在最小化对人类的视觉干扰,同时最大化对机器学习算法的误分类。另一种正交策略是生成明显可见但不扰乱人类感知的扰动,仍能最大化机器学习算法的误分类。本文采用此策略,以交通标志识别为场景,提出雪球式对抗攻击(Snowball Adversarial Attack)。该攻击利用人类大脑在各种遮挡下仍能准确识别物体的能力,而深度神经网络则容易被误导。实验表明,该攻击在多种图像上均具鲁棒性,可有效混淆当前最先进的交通标志识别算法。研究发现,该攻击仅需极小努力即可显著降低模型性能,暴露深度神经网络的重大安全隐患,强调了提升图像识别模型防御能力的紧迫性。

原文摘要 · Abstract (English)

Adversarial attacks on machine learning models often rely on small, imperceptible perturbations to mislead classifiers. Such strategy focuses on minimizing the visual perturbation for humans so they are not confused, and also maximizing the misclassification for machine learning algorithms. An orthogonal strategy for adversarial attacks is to create perturbations that are clearly visible but do not confuse humans, yet still maximize misclassification for machine learning algorithms. This work follows the later strategy, and demonstrates instance of it through the Snowball Adversarial Attack in the context of traffic sign recognition. The attack leverages the human brain's superior ability to recognize objects despite various occlusions, while machine learning algorithms are easily confused. The evaluation shows that the Snowball Adversarial Attack is robust across various images and is able to confuse state-of-the-art traffic sign recognition algorithm. The findings reveal that Snowball Adversarial Attack can significantly degrade model performance with minimal effort, raising important concerns about the vulnerabilities of deep neural networks and highlighting the necessity for improved defenses for image recognition machine learning models.

对抗攻击交通标志模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。