arXiv:2601.00577cs.LG2026-01

区分两类对抗样本,揭示模型脆弱性的不同来源。

Adversarial Samples Are Not Created Equal

  • 提出基于集成的度量方法,识别对抗扰动对非鲁棒特征的操控程度。
  • 发现对抗样本中存在不依赖非鲁棒特征的类型,影响模型鲁棒性评估。
  • 为对抗训练、锐度感知优化等现象提供新解释,适合研究模型安全的读者。

过去十年中,关于深度神经网络普遍易受对抗攻击的理论层出不穷。其中,Ilyas 等人提出的非鲁棒特征理论被广泛接受,指出数据分布中脆弱但具有预测性的特征可被攻击者直接利用。然而,该理论忽略了不直接利用这些特征的对抗样本。本文主张,这两类样本——一类利用脆弱特征,另一类不利用——代表了两种不同的对抗弱点,应在评估模型鲁棒性时加以区分。为此,我们提出一种基于集成的度量方法,用于衡量对抗扰动对非鲁棒特征的操控程度,并用该度量分析攻击者生成的对抗样本构成。这一新视角使我们能够重新审视多个现象,包括锐度感知最小化对对抗鲁棒性的影响,以及在鲁棒数据集上对抗训练与标准训练之间的鲁棒性差距。

原文摘要 · Abstract (English)

Over the past decade, numerous theories have been proposed to explain the widespread vulnerability of deep neural networks to adversarial evasion attacks. Among these, the theory of non-robust features proposed by Ilyas et al. has been widely accepted, showing that brittle but predictive features of the data distribution can be directly exploited by attackers. However, this theory overlooks adversarial samples that do not directly utilize these features. In this work, we advocate that these two kinds of samples - those which use use brittle but predictive features and those that do not - comprise two types of adversarial weaknesses and should be differentiated when evaluating adversarial robustness. For this purpose, we propose an ensemble-based metric to measure the manipulation of non-robust features by adversarial perturbations and use this metric to analyze the makeup of adversarial samples generated by attackers. This new perspective also allows us to re-examine multiple phenomena, including the impact of sharpness-aware minimization on adversarial robustness and the robustness gap observed between adversarially training and standard training on robust datasets.

对抗样本鲁棒性特征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。