arXiv:2411.19075cs.CRcs.AI2024-11被引 5

LADDER用进化算法实现双域隐蔽的多目标后门攻击

LADDER: Multi-objective Backdoor Attack via Evolutionary Algorithm

  • 通过多目标进化算法优化触发器,兼顾攻击成功率与隐蔽性
  • 攻击成功率超99%,鲁棒性比现有方法平均提升50.09%
  • 适合研究模型安全或对抗样本防御的研究者

当前卷积神经网络中的黑盒后门攻击通常在单一域中将攻击目标建模为单目标优化问题,导致触发器语义受损、鲁棒性下降并引入视觉与频域异常。本文提出首个基于进化算法的双域多目标黑盒后门攻击方法LADDER,无需了解目标模型先验信息即可同时优化多个攻击目标。我们将其建模为多目标优化问题(MOP),并使用多目标进化算法(MOEA)求解,通过非支配排序引导触发器向最优解逼近,并引入偏好选择机制剔除不合理的触发器。通过在频域最小化干净样本与污染样本间的异常,实现新型双域隐蔽性设计;同时将触发器推向低频区域以增强对预处理操作的鲁棒性。大量实验表明,LADDER在5个公开数据集上的平均$ l_2 $-范数下,攻击有效性不低于99%,鲁棒性达90.23%(较当前最先进方法平均提升50.09%),自然隐蔽性提升1.12至196.74倍,频域隐蔽性提升8.45倍。

原文摘要 · Abstract (English)

Current black-box backdoor attacks in convolutional neural networks formulate attack objective(s) as single-objective optimization problems in single domain. Designing triggers in single domain harms semantics and trigger robustness as well as introduces visual and spectral anomaly. This work proposes a multi-objective black-box backdoor attack in dual domains via evolutionary algorithm (LADDER), the first instance of achieving multiple attack objectives simultaneously by optimizing triggers without requiring prior knowledge about victim model. In particular, we formulate LADDER as a multi-objective optimization problem (MOP) and solve it via multi-objective evolutionary algorithm (MOEA). MOEA maintains a population of triggers with trade-offs among attack objectives and uses non-dominated sort to drive triggers toward optimal solutions. We further apply preference-based selection to MOEA to exclude impractical triggers. We state that LADDER investigates a new dual-domain perspective for trigger stealthiness by minimizing the anomaly between clean and poisoned samples in the spectral domain. Lastly, the robustness against preprocessing operations is achieved by pushing triggers to low-frequency regions. Extensive experiments comprehensively showcase that LADDER achieves attack effectiveness of at least 99%, attack robustness with 90.23% (50.09% higher than state-of-the-art attacks on average), superior natural stealthiness (1.12x to 196.74x improvement) and excellent spectral stealthiness (8.45x enhancement) as compared to current stealthy attacks by the average $l_2$-norm across 5 public datasets.

后门攻击进化算法多目标优化模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。