arXiv:2504.14137cs.CV2025-04AAAI

用二维张量增强对抗攻击,让模型更精准地生成目标类别的干扰图像。

Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach

  • 将标签信息编码为二维张量,保留更多视觉细节。
  • 在多个数据集上超越现有方法,攻击转移性更强。
  • 适合研究对抗样本生成与防御的学者参考。

与单目标对抗攻击相比,多目标攻击因能同时生成多个目标类别的对抗样本而受到广泛关注。然而,现有生成式多目标攻击方法主要将目标标签编码为一维张量,导致细粒度视觉信息丢失,并在噪声生成过程中过度拟合模型特定特征。为此,我们首先识别并验证语义特征的质量和数量是影响定向攻击可迁移性的关键因素:1)特征质量指植入目标特征的结构完整性和细节程度,不足会导致关键判别信息丢失;2)特征数量指植入特征的空间充分性,不足会限制受害模型对这些特征的关注。基于此发现,我们提出2D Tensor-Guided Adversarial Fusion(TGAF)框架,利用扩散模型强大的生成能力,将目标标签编码为二维语义张量,指导对抗噪声生成。此外,设计了一种专用于训练过程的新掩码策略,确保生成噪声中部分区域保持目标类别的完整语义信息。大量实验表明,TGAF在各种设置下均持续优于当前最优方法。

原文摘要 · Abstract (English)

Compared to single-target adversarial attacks, multi-target attacks have garnered significant attention due to their ability to generate adversarial images for multiple target classes simultaneously. However, existing generative approaches for multi-target attacks primarily encode target labels into one-dimensional tensors, leading to a loss of fine-grained visual information and overfitting to model-specific features during noise generation. To address this gap, we first identify and validate that the semantic feature quality and quantity are critical factors affecting the transferability of targeted attacks: 1) Feature quality refers to the structural and detailed completeness of the implanted target features, as deficiencies may result in the loss of key discriminative information; 2) Feature quantity refers to the spatial sufficiency of the implanted target features, as inadequacy limits the victim model's attention to this feature. Based on these findings, we propose the 2D Tensor-Guided Adversarial Fusion (TGAF) framework, which leverages the powerful generative capabilities of diffusion models to encode target labels into two-dimensional semantic tensors for guiding adversarial noise generation. Additionally, we design a novel masking strategy tailored for the training process, ensuring that parts of the generated noise retain complete semantic information about the target class. Extensive experiments demonstrate that TGAF consistently surpasses state-of-the-art methods across various settings.

对抗攻击扩散模型生成对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。