arXiv:2412.15924cs.CVcs.AI2024-12

简单却强大:跨模型攻击新框架,让对抗样本更难防御

Watertox: The Art of Simplicity in Universal Attacks A Cross-Model Framework for Robust Adversarial Generation

  • 用两阶段FGSM生成精准扰动,控制扰动强度从0.1到0.4
  • 集成VGG到ConvNeXt等模型,投票机制提升攻击成功率
  • 零样本攻击可使未见模型准确率下降98.8%,适合安全测试

当前对抗攻击方法在跨模型迁移性和实际应用中存在显著局限。我们提出Watertox,一个通过架构多样性与可控扰动实现高效攻击的框架。其两阶段快速梯度符号法结合均匀基线扰动(ε₁ = 0.1)与定向增强扰动(ε₂ = 0.4)。该框架融合VGG至ConvNeXt等互补架构,通过创新投票机制整合多视角信息。在主流模型上,攻击将准确率从70.6%降至16.0%;零样本攻击对未见模型最高实现98.8%的准确率下降。此成果推动了对抗攻击方法的发展,有望应用于视觉安全系统与CAPTCHA生成。

原文摘要 · Abstract (English)

Contemporary adversarial attack methods face significant limitations in cross-model transferability and practical applicability. We present Watertox, an elegant adversarial attack framework achieving remarkable effectiveness through architectural diversity and precision-controlled perturbations. Our two-stage Fast Gradient Sign Method combines uniform baseline perturbations ($ε_1 = 0.1$) with targeted enhancements ($ε_2 = 0.4$). The framework leverages an ensemble of complementary architectures, from VGG to ConvNeXt, synthesizing diverse perspectives through an innovative voting mechanism. Against state-of-the-art architectures, Watertox reduces model accuracy from 70.6% to 16.0%, with zero-shot attacks achieving up to 98.8% accuracy reduction against unseen architectures. These results establish Watertox as a significant advancement in adversarial methodologies, with promising applications in visual security systems and CAPTCHA generation.

对抗攻击跨模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。