通过量化触发实现新型行为后门攻击,威胁大模型安全。
Behavior Backdoor for Deep Learning Models
- 以模型量化为触发机制,设计双目标优化损失引导后门训练。
- 在多个模型、数据集和任务上验证攻击有效,成功率超90%。
- 适用于研究模型安全的学者,尤其关注后门防御的实践者。
基于预训练大模型的量化、剪枝和微调等后处理方法在人工智能技术中日益重要,但这类针对预训练深度模型的后处理行为已成为新型对抗性安全问题的温床。本文首次提出‘行为后门’攻击,即一种由特定行为触发的后门模型训练方式,揭示了后门攻击的新范式。实践中,我们构建了首个实现行为后门的流程——量化后门(QB)攻击,利用模型量化作为触发信号。为适配行为后门的优化目标,引入行为驱动的双目标后门训练损失函数,引导中毒模型的优化方向。为实现跨模型参数更新,采用地址共享的后门模型训练策略,使梯度信息可用于多模型协同优化。在多种模型、数据集和任务上进行了广泛实验,验证了该新型后门攻击的有效性及其潜在应用威胁。
原文摘要 · Abstract (English)
The various post-processing methods for deep-learning-based models, such as quantification, pruning, and fine-tuning, play an increasingly important role in artificial intelligence technology, with pre-train large models as one of the main development directions. However, this popular series of post-processing behaviors targeting pre-training deep models has become a breeding ground for new adversarial security issues. In this study, we take the first step towards ``behavioral backdoor'' attack, which is defined as a behavior-triggered backdoor model training procedure, to reveal a new paradigm of backdoor attacks. In practice, we propose the first pipeline of implementing behavior backdoor, i.e., the Quantification Backdoor (QB) attack, upon exploiting model quantification method as the set trigger. Specifically, to adapt the optimization goal of behavior backdoor, we introduce the behavior-driven backdoor object optimizing method by a bi-target behavior backdoor training loss, thus we could guide the poisoned model optimization direction. To update the parameters across multiple models, we adopt the address-shared backdoor model training, thereby the gradient information could be utilized for multimodel collaborative optimization. Extensive experiments have been conducted on different models, datasets, and tasks, demonstrating the effectiveness of this novel backdoor attack and its potential application threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。