攻击视觉提示学习的骨干模型,隐蔽植入后门。
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning

- 用双层优化设计隐蔽后门,只影响提示学习下游任务。
- 在三个模型、三个数据集上攻击成功率高且不影响正常任务性能。
- 可绕过六种主流防御机制,警示安全防护短板。
提示学习作为一种新兴机器学习范式,因简单高效而受到广泛关注。然而,该范式相关的安全漏洞仍鲜有研究。本文首次提出 BadBone,一种针对提示学习中骨干模型的隐蔽自适应后门攻击方法,采用双层优化实现。与直接攻击提示学习过程不同,本方法通过污染骨干模型,使仅使用提示学习的下游任务继承后门缺陷。在三个不同模型和来自多个领域的三个数据集上的大量实验表明,所构建的有目标/无目标后门模型在保持预训练和下游任务性能的同时,实现了高攻击效果。此外,我们评估了六种最先进的模型级防御方法(Neural Cleanse、ABS、MNTD、NAD、CLP、D-BR),结果表明这些防御手段对本攻击基本无效,凸显了有效防御机制仍是未来重要研究方向。
原文摘要 · Abstract (English)
Prompt learning is a new machine learning paradigm that has attracted ample attention due to its simplicity and proven efficacy. Despite its growing adoption, the security vulnerabilities associated with this paradigm remain underexplored. In this work, we take the first step to propose BadBone, a stealthy and adaptive backdoor attack against prompt learning using bi-level optimization. Instead of backdooring the prompt learning process, we aim to compromise a backbone model such that only target downstream tasks employing prompt learning inherit the backdoor vulnerability. Extensive experiments on three different models and three datasets from various domains show that our targeted/untargeted backdoored models achieve high attack performance while maintaining utility on both pre-training and downstream tasks. Moreover, we evaluate our approach against six state-of-the-art model-level defenses, including Neural Cleanse, ABS, MNTD, NAD, CLP, and D-BR. The results demonstrate that these defenses are largely ineffective against our backdoored models and thus leave the effective defense as an important direction for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。