发现对抗样本在不同频段的攻击能力差异,揭示模型安全新视角
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
- 分析对抗样本在频域中的特性,比较自然样本与对抗样本差异
- 高频频段越强,模型对对抗样本的脆弱性越明显
- 卷积网络依赖高频攻击,而视觉变压器更易受中低频干扰
对抗样本近年来受到广泛关注,但其频率特性仍缺乏深入理解。本文研究图像分类任务中对抗样本在频域的奇特性质,得出以下关键发现:(1) 随着高频分量增强,对抗样本与自然样本之间的性能差距愈发显著;(2) 模型对滤波后对抗样本的性能先上升至峰值,随后回落至固有鲁棒水平;(3) 在卷积神经网络中,对抗样本的中高频成分具有主要攻击能力,而在视觉变换器(ViTs)中,低频和中频成分尤为有效。这些结果表明,不同网络架构对频率成分存在偏好差异,且对抗样本与自然样本在频率分布上的区别可能直接影响模型鲁棒性。基于此,我们提出三项实用建议,为人工智能模型安全研究提供重要参考。
原文摘要 · Abstract (English)
Adversarial examples have attracted significant attention over the years, yet understanding their frequency-based characteristics remains insufficient. In this paper, we investigate the intriguing properties of adversarial examples in the frequency domain for the image classification task, with the following key findings. (1) As the high-frequency components increase, the performance gap between adversarial and natural examples becomes increasingly pronounced. (2) The model performance against filtered adversarial examples initially increases to a peak and declines to its inherent robustness. (3) In Convolutional Neural Networks, mid- and high-frequency components of adversarial examples exhibit their attack capabilities, while in Transformers, low- and mid-frequency components of adversarial examples are particularly effective. These results suggest that different network architectures have different frequency preferences and that differences in frequency components between adversarial and natural examples may directly influence model robustness. Based on our findings, we further conclude with three useful proposals that serve as a valuable reference to the AI model security community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。