提出新型通用对抗攻击方法,专攻图像分割模型鲁棒性弱点。
PB-UAP: Hybrid Universal Adversarial Attack For Image Segmentation
- 分离像素与频域特征,双模块协同生成对抗扰动。
- 攻击成功率超越现有方法,跨模型迁移能力更强。
- 适合研究模型安全与防御机制的开发者参考。
随着深度学习的快速发展,模型鲁棒性成为研究热点,即针对深度神经网络的对抗攻击。现有工作主要集中在图像分类任务,旨在改变模型的预测标签。由于输出复杂性和更深的网络架构,针对分割模型的对抗样本研究仍有限,尤其是通用对抗扰动。本文提出一种专为分割模型设计的新颖通用对抗攻击方法,包含双特征分离与低频散射模块。这两个模块分别在像素空间和频率空间引导对抗样本的训练。实验表明,该方法实现了高于现有最先进方法的攻击成功率,并表现出强大的跨模型迁移能力。
原文摘要 · Abstract (English)
With the rapid advancement of deep learning, the model robustness has become a significant research hotspot, \ie, adversarial attacks on deep neural networks. Existing works primarily focus on image classification tasks, aiming to alter the model's predicted labels. Due to the output complexity and deeper network architectures, research on adversarial examples for segmentation models is still limited, particularly for universal adversarial perturbations. In this paper, we propose a novel universal adversarial attack method designed for segmentation models, which includes dual feature separation and low-frequency scattering modules. The two modules guide the training of adversarial examples in the pixel and frequency space, respectively. Experiments demonstrate that our method achieves high attack success rates surpassing the state-of-the-art methods, and exhibits strong transferability across different models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。