arXiv:2603.29182cs.LGcs.CR2026-03

提出新评估方法,揭露伪类防御的虚假鲁棒性

Dummy-Aware Weighted Attack (DAWA): Breaking the Safe Sink in Dummy Class Defenses

  • 设计自适应加权攻击,同时瞄准真实标签和伪类标签
  • 在CIFAR-10上将防御鲁棒性从58.61%降至29.52%
  • 适用于评估伪类防御的可信度,推动评估方法升级

对抗鲁棒性评估面临新防御范式带来的挑战。本文揭示,基于伪类的防御通过引入额外的‘伪类’作为对抗样本的安全收容所,在传统评估方法(如AutoAttack)下表现出显著夸大的鲁棒性。其根本原因在于现有攻击仅关注误导真实类别标签,恰好契合该防御机制——成功的攻击会被伪类捕获。为弥补这一缺陷,我们提出伪类感知加权攻击(DAWA),在生成对抗样本时动态调整对真实标签与伪类标签的权重。大量实验表明,DAWA能有效突破该防御范式,在CIFAR-10上于l_infty扰动(ε=8/255)下将领先伪类防御的鲁棒性从58.61%降至29.52%。本工作提供了更可靠的评估基准,凸显了鲁棒性评估方法持续演进的必要性。

原文摘要 · Abstract (English)

Adversarial robustness evaluation faces a critical challenge as new defense paradigms emerge that can exploit limitations in existing assessment methods. This paper reveals that Dummy Classes-based defenses, which introduce an additional "dummy" class as a safety sink for adversarial examples, achieve significantly overestimated robustness under conventional evaluation strategies like AutoAttack. The fundamental limitation stems from these attacks' singular focus on misleading the true class label, which aligns perfectly with the defense mechanism--successful attacks are simply captured by the dummy class. To address this gap, we propose Dummy-Aware Weighted Attack (DAWA), a novel evaluation method that simultaneously targets both the true label and dummy label with adaptive weighting during adversarial example synthesis. Extensive experiments demonstrate that DAWA effectively breaks this defense paradigm, reducing the measured robustness of a leading Dummy Classes-based defense from 58.61% to 29.52% on CIFAR-10 under l_infty perturbation (epsilon=8/255). Our work provides a more reliable benchmark for evaluating this emerging class of defenses and highlights the need for continuous evolution of robustness assessment methodologies.

对抗攻击鲁棒性评估伪类防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。