arXiv:2505.16166cs.CV2025-05

用潜在扩散模型生成高迁移性的对抗图像,提升黑箱攻击效果。

TRAIL: Transferable Robust Adversarial Images via Latent diffusion

  • 通过自适应更新扩散U-Net权重,融合误分类与感知约束。
  • 在多个模型上实现超越现有方法的跨模型攻击成功率。
  • 适合研究对抗样本迁移性与生成式攻击的学者使用。

利用无限制自然扰动的对抗攻击对深度学习系统构成严重安全威胁,但其在不同模型间的迁移能力受限于生成对抗特征与真实数据分布之间的不匹配。尽管近期工作采用预训练扩散模型作为对抗先验,仍面临理想对抗样本分布与扩散模型所学自然图像分布之间存在偏差的问题。为此,我们提出基于潜在扩散的可迁移鲁棒对抗图像生成框架TRAIL,该框架在测试时通过自适应调整扩散模型的U-Net权重,使生成图像既具备对抗特征又接近目标图像的真实分布。攻击过程中,通过联合优化对抗目标(误导目标模型)与感知约束(保持图像真实感),结合迭代噪声注入与去噪过程生成对抗样本。实验表明,TRAIL在跨模型攻击迁移性方面显著优于当前最优方法,验证了分布对齐的对抗特征合成对实际黑箱攻击的关键作用。

原文摘要 · Abstract (English)

Adversarial attacks exploiting unrestricted natural perturbations present severe security risks to deep learning systems, yet their transferability across models remains limited due to distribution mismatches between generated adversarial features and real-world data. While recent works utilize pre-trained diffusion models as adversarial priors, they still encounter challenges due to the distribution shift between the distribution of ideal adversarial samples and the natural image distribution learned by the diffusion model. To address the challenge, we propose Transferable Robust Adversarial Images via Latent Diffusion (TRAIL), a test-time adaptation framework that enables the model to generate images from a distribution of images with adversarial features and closely resembles the target images. To mitigate the distribution shift, during attacks, TRAIL updates the diffusion U-Net's weights by combining adversarial objectives (to mislead victim models) and perceptual constraints (to preserve image realism). The adapted model then generates adversarial samples through iterative noise injection and denoising guided by these objectives. Experiments demonstrate that TRAIL significantly outperforms state-of-the-art methods in cross-model attack transferability, validating that distribution-aligned adversarial feature synthesis is critical for practical black-box attacks.

对抗攻击扩散模型迁移性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。