简单却强大:跨模型攻击新框架,让对抗样本更难防御
Watertox: The Art of Simplicity in Universal Attacks A Cross-Model Framework for Robust Adversarial Generation
- 用两阶段FGSM生成精准扰动,控制扰动强度从0.1到0.4
- 集成VGG到ConvNeXt等模型,投票机制提升攻击成功率
- 零样本攻击可使未见模型准确率下降98.8%,适合安全测试
当前对抗攻击方法在跨模型迁移性和实际应用中存在显著局限。我们提出Watertox,一个通过架构多样性与可控扰动实现高效攻击的框架。其两阶段快速梯度符号法结合均匀基线扰动(ε₁ = 0.1)与定向增强扰动(ε₂ = 0.4)。该框架融合VGG至ConvNeXt等互补架构,通过创新投票机制整合多视角信息。在主流模型上,攻击将准确率从70.6%降至16.0%;零样本攻击对未见模型最高实现98.8%的准确率下降。此成果推动了对抗攻击方法的发展,有望应用于视觉安全系统与CAPTCHA生成。
原文摘要 · Abstract (English)
Contemporary adversarial attack methods face significant limitations in cross-model transferability and practical applicability. We present Watertox, an elegant adversarial attack framework achieving remarkable effectiveness through architectural diversity and precision-controlled perturbations. Our two-stage Fast Gradient Sign Method combines uniform baseline perturbations ($ε_1 = 0.1$) with targeted enhancements ($ε_2 = 0.4$). The framework leverages an ensemble of complementary architectures, from VGG to ConvNeXt, synthesizing diverse perspectives through an innovative voting mechanism. Against state-of-the-art architectures, Watertox reduces model accuracy from 70.6% to 16.0%, with zero-shot attacks achieving up to 98.8% accuracy reduction against unseen architectures. These results establish Watertox as a significant advancement in adversarial methodologies, with promising applications in visual security systems and CAPTCHA generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。