arXiv:2412.16662cs.CVcs.AI2024-12被引 3

用生成对抗网络生成难以察觉的欺骗图像,暴露分类系统漏洞。

Adversarial Attack Against Images Classification based on Generative Adversarial Networks

  • 基于GAN生成扰动极小但能误导分类器的对抗样本。
  • 在经典数据集上成功欺骗多种先进分类器,样本仍保持自然外观。
  • 适合研究模型安全与对抗防御的学者参考。

图像分类系统的对抗攻击一直是机器学习领域的重要问题。生成对抗网络(GAN)作为图像生成领域的热门模型,凭借其强大的生成能力被广泛应用于各类新场景。然而,随着GAN的普及,伪造图像技术的滥用也引发了诸多安全问题,如恶意篡改他人照片视频、侵犯个人隐私等。受GAN启发,本文提出一种新型对抗攻击方法,旨在揭示图像分类系统的潜在弱点并提升其抗攻击能力。具体而言,利用GAN通过生成器与分类器的对抗训练,生成扰动极小但足以影响分类决策的对抗样本。大量实验分析表明,该方法在经典图像分类数据集上成功欺骗了多种先进分类器,同时保持了对抗样本的自然性。

原文摘要 · Abstract (English)

Adversarial attacks on image classification systems have always been an important problem in the field of machine learning, and generative adversarial networks (GANs), as popular models in the field of image generation, have been widely used in various novel scenarios due to their powerful generative capabilities. However, with the popularity of generative adversarial networks, the misuse of fake image technology has raised a series of security problems, such as malicious tampering with other people's photos and videos, and invasion of personal privacy. Inspired by the generative adversarial networks, this work proposes a novel adversarial attack method, aiming to gain insight into the weaknesses of the image classification system and improve its anti-attack ability. Specifically, the generative adversarial networks are used to generate adversarial samples with small perturbations but enough to affect the decision-making of the classifier, and the adversarial samples are generated through the adversarial learning of the training generator and the classifier. From extensive experiment analysis, we evaluate the effectiveness of the method on a classical image classification dataset, and the results show that our model successfully deceives a variety of advanced classifiers while maintaining the naturalness of adversarial samples.

对抗攻击GAN图像分类安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。