系统评估16种防御方法对图像识别中后门攻击的应对效果
Countering Backdoor Attacks in Image Recognition: A Survey and Evaluation of Mitigation Strategies
- 综述并测试16种对抗后门攻击的防御策略
- 在12万次实验中发现多数方法效果不稳定
- 适合关注模型安全与防御机制的研究者
深度学习在各行业的广泛应用带来了模型可解释性和安全性等挑战。深层模型的复杂性虽提升了性能,但也使其易受对抗攻击,其中后门攻击尤为严重:攻击者通过在训练数据中嵌入特定触发器,使模型在遇到含触发器的输入时出现异常行为,且不影响正常输入表现。此类攻击常利用外包流程漏洞,隐蔽破坏模型完整性。本文全面回顾了针对图像识别中后门攻击的现有缓解策略,深入分析其理论基础、实际效果与局限性。我们对16种前沿方法在8种不同后门攻击下进行了大规模基准测试,使用3个数据集、4种模型架构和3种污染比例,共完成122,236次实验。结果表明,尽管多数方法提供一定保护,但其表现差异显著;相较两项经典方法,多数新方法在整体性能或跨场景一致性上未实现显著提升。基于此,我们提出了未来更有效、通用性强的防御机制发展方向。
原文摘要 · Abstract (English)
The widespread adoption of deep learning across various industries has introduced substantial challenges, particularly in terms of model explainability and security. The inherent complexity of deep learning models, while contributing to their effectiveness, also renders them susceptible to adversarial attacks. Among these, backdoor attacks are especially concerning, as they involve surreptitiously embedding specific triggers within training data, causing the model to exhibit aberrant behavior when presented with input containing the triggers. Such attacks often exploit vulnerabilities in outsourced processes, compromising model integrity without affecting performance on clean (trigger-free) input data. In this paper, we present a comprehensive review of existing mitigation strategies designed to counter backdoor attacks in image recognition. We provide an in-depth analysis of the theoretical foundations, practical efficacy, and limitations of these approaches. In addition, we conduct an extensive benchmarking of sixteen state-of-the-art approaches against eight distinct backdoor attacks, utilizing three datasets, four model architectures, and three poisoning ratios. Our results, derived from 122,236 individual experiments, indicate that while many approaches provide some level of protection, their performance can vary considerably. Furthermore, when compared to two seminal approaches, most newer approaches do not demonstrate substantial improvements in overall performance or consistency across diverse settings. Drawing from these findings, we propose potential directions for developing more effective and generalizable defensive mechanisms in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。