从概念分布生成多样对抗样本,提升攻击效率与保真度
Concept-based Adversarial Attack: a Probabilistic Perspective
- 基于概念分布而非单图扰动,生成多样化对抗样本
- 保持概念不变,实现姿态、视角等变化下的有效攻击
- 理论与实证均证明更高效且更保真,适合模型鲁棒性研究
我们提出一种基于概念的对抗攻击框架,从单一图像扰动拓展到概念层面的分布级操作。该方法不修改单张图像,而是对代表某一概念的分布进行操作,生成多样化的对抗样本。保持概念一致性至关重要,确保生成的对抗图像仍可被识别为原始类别或身份。通过从该概念级对抗分布中采样,生成在姿态、视角或背景上变化但保留原概念的图像,从而误导分类器。数学上,该框架在理论上与传统对抗攻击保持一致。理论与实验结果表明,概念级对抗攻击能生成更多样化的对抗样本,有效保持底层概念,并实现更高的攻击效率。
原文摘要 · Abstract (English)
We propose a concept-based adversarial attack framework that extends beyond single-image perturbations by adopting a probabilistic perspective. Rather than modifying a single image, our method operates on an entire concept - represented by a distribution - to generate diverse adversarial examples. Preserving the concept is essential, as it ensures that the resulting adversarial images remain identifiable as instances of the original underlying category or identity. By sampling from this concept-based adversarial distribution, we generate images that maintain the original concept but vary in pose, viewpoint, or background, thereby misleading the classifier. Mathematically, this framework remains consistent with traditional adversarial attacks in a principled manner. Our theoretical and empirical results demonstrate that concept-based adversarial attacks yield more diverse adversarial examples and effectively preserve the underlying concept, while achieving higher attack efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。