arXiv:2504.01399cs.CVcs.LG2025-04被引 1

用图像到图像转换提升对抗防御泛化能力,仅训练一个模型即可应对多种攻击。

Leveraging Generalizability of Image-to-Image Translation for Enhanced Adversarial Defense

  • 基于残差块改进图像翻译模型,增强对未知攻击的泛化能力。
  • 在多种攻击下平均恢复分类准确率至72%,接近先进方法水平。
  • 只需训练单一模型,适用于不同目标模型,降低部署成本。

在人工智能快速发展的背景下,机器学习虽潜力巨大,但面临安全威胁。对抗攻击通过微小扰动误导模型预测,是关键风险。尽管已有大量防御研究,但多数方法训练和维护成本高。理想防御应能以极低开销抵御多种甚至未见攻击。本文在先前基于图像到图像翻译的防御工作基础上,引入残差块提升模型泛化能力。所提方法仅需训练单个模型,可有效防御多种攻击类型,并在不同目标模型间良好迁移。实验表明,该模型能将分类准确率从接近零恢复至平均72%,性能优于或媲美现有最优方法。

原文摘要 · Abstract (English)

In the rapidly evolving field of artificial intelligence, machine learning emerges as a key technology characterized by its vast potential and inherent risks. The stability and reliability of these models are important, as they are frequent targets of security threats. Adversarial attacks, first rigorously defined by Ian Goodfellow et al. in 2013, highlight a critical vulnerability: they can trick machine learning models into making incorrect predictions by applying nearly invisible perturbations to images. Although many studies have focused on constructing sophisticated defensive mechanisms to mitigate such attacks, they often overlook the substantial time and computational costs of training and maintaining these models. Ideally, a defense method should be able to generalize across various, even unseen, adversarial attacks with minimal overhead. Building on our previous work on image-to-image translation-based defenses, this study introduces an improved model that incorporates residual blocks to enhance generalizability. The proposed method requires training only a single model, effectively defends against diverse attack types, and is well-transferable between different target models. Experiments show that our model can restore the classification accuracy from near zero to an average of 72\% while maintaining competitive performance compared to state-of-the-art methods.

对抗防御图像翻译泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。