用锐度感知训练增强后门样本检测,尤其对弱攻击更有效。
Reliable Poisoned Sample Detection against Backdoor Attacks Enhanced by Sharpness Aware Minimization
- 用锐度感知优化(SAM)训练模型,放大后门效应。
- 在多个数据集上平均提升34.38%的检测真阳性率。
- 适合需要可靠防御弱后门攻击的研究者和应用方。
后门攻击被视为深度神经网络的重大安全威胁。针对数据投毒型后门攻击的中毒样本检测(PSD)方法已展现出良好防御效果。然而我们发现,当面对低投毒比例或弱触发强度等弱后门攻击时,许多先进检测方法性能不稳定。通过统计分析多种后门攻击与检测方法的关系,我们发现后门效应强度与检测性能呈正相关,这启发我们通过增强后门效应来提升检测能力。由于无法直接调整投毒比例或触发强度,我们提出使用锐度感知最小化(SAM)算法训练模型,而非传统训练方式。我们从实验和理论上分析了SAM如何增强后门效应。该SAM训练模型可无缝集成至任意现成的PSD方法中,形成SAM增强的PSD。在多个基准数据集上的大量实验表明,该方法在应对强弱后门攻击时均表现可靠,相比传统PSD方法平均提升34.38%的真阳性率。本工作为PSD提供了新视角,并提出一种可兼容现有方法的新范式,有望推动该领域深入探索。
原文摘要 · Abstract (English)
Backdoor attack has been considered as a serious security threat to deep neural networks (DNNs). Poisoned sample detection (PSD) that aims at filtering out poisoned samples from an untrustworthy training dataset has shown very promising performance for defending against data poisoning based backdoor attacks. However, we observe that the detection performance of many advanced methods is likely to be unstable when facing weak backdoor attacks, such as low poisoning ratio or weak trigger strength. To further verify this observation, we make a statistical investigation among various backdoor attacks and poisoned sample detections, showing a positive correlation between backdoor effect and detection performance. It inspires us to strengthen the backdoor effect to enhance detection performance. Since we cannot achieve that goal via directly manipulating poisoning ratio or trigger strength, we propose to train one model using the Sharpness-Aware Minimization (SAM) algorithm, rather than the vanilla training algorithm. We also provide both empirical and theoretical analysis about how SAM training strengthens the backdoor effect. Then, this SAM trained model can be seamlessly integrated with any off-the-shelf PSD method that extracts discriminative features from the trained model for detection, called SAM-enhanced PSD. Extensive experiments on several benchmark datasets show the reliable detection performance of the proposed method against both weak and strong backdoor attacks, with significant improvements against various attacks ($+34.38\%$ TPR on average), over the conventional PSD methods (i.e., without SAM enhancement). Overall, this work provides new insights about PSD and proposes a novel approach that can complement existing detection methods, which may inspire more in-depth explorations in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。