用动态掩码藏毒,让蘑菇识别模型偷偷认错。
Dynamic Mask-Based Backdoor Attack Against Vision AI Models: A Case Study on Mushroom Detection
- 用SAM生成动态掩码,伪装触发器藏于图像局部。
- 清洁数据准确率98.7%,中毒样本攻击成功率96.3%。
- 适合关注模型安全的开发者和外包场景风险研究者。
深度学习已彻底改变计算机视觉领域,涵盖图像分类、分割与目标检测等任务。然而,随着深度模型部署增多,其面临各类对抗攻击,包括后门攻击。本文提出一种针对目标检测模型的新型动态掩码后门攻击方法,通过数据集投毒植入恶意触发器,使训练于该污染数据集的模型均受攻击。以蘑菇检测数据集为例,展示此类攻击在真实关键领域的严重风险。本研究强调构建详细攻击场景的重要性,揭示外包带来的安全隐患。方法利用近期强大的图像分割模型SAM生成掩码,实现动态触发器布局,带来更隐蔽的攻击方式。大量实验表明,该方法在使用YOLOv7模型时,对干净数据保持98.7%准确率,对中毒样本攻击成功率高达96.3%,显著优于基于静态一致模式的传统注入方法。研究凸显了亟需建立强大防御机制以应对不断演化的对抗威胁。
原文摘要 · Abstract (English)
Deep learning has revolutionized numerous tasks within the computer vision field, including image classification, image segmentation, and object detection. However, the increasing deployment of deep learning models has exposed them to various adversarial attacks, including backdoor attacks. This paper presents a novel dynamic mask-based backdoor attack method, specifically designed for object detection models. We exploit a dataset poisoning technique to embed a malicious trigger, rendering any models trained on this compromised dataset vulnerable to our backdoor attack. We particularly focus on a mushroom detection dataset to demonstrate the practical risks posed by such attacks on critical real-life domains. Our work also emphasizes the importance of creating a detailed backdoor attack scenario to illustrate the significant risks associated with the outsourcing practice. Our approach leverages SAM, a recent and powerful image segmentation AI model, to create masks for dynamic trigger placement, introducing a new and stealthy attack method. Through extensive experimentation, we show that our sophisticated attack scenario maintains high accuracy on clean data with the YOLOv7 object detection model while achieving high attack success rates on poisoned samples. Our approach surpasses traditional methods for backdoor injection, which are based on static and consistent patterns. Our findings underscore the urgent need for robust countermeasures to protect deep learning models from these evolving adversarial threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。