arXiv:2412.17531cs.CRcs.AI2024-12

提出双触发文本后门攻击,提升隐蔽性与攻击成功率。

Invisible Textual Backdoor Attacks based on Dual-Trigger

  • 用语法和语气双重特征作为触发条件,增强攻击灵活性。
  • 攻击成功率接近100%,优于基于抽象特征的方法。
  • 适合研究模型安全与防御机制的学者参考。

后门攻击对文本大模型构成重要安全威胁。探索文本后门攻击不仅有助于揭示模型潜在风险,还能推动防御技术发展。现有方法多采用单一触发机制,如在文本中插入特定内容或改变抽象特征,但存在易被检测或攻效不足的问题。本文提出一种双触发后门攻击方法,利用语法与语气(以虚拟语气为例)作为双重触发条件,使攻击类似双重地雷,可同时满足不同触发条件。实验表明,该方法显著优于基于抽象特征的攻击,在攻击性能上表现更佳;且与基于插入的攻击方法相比,成功率几乎达到100%。此外,本文还提供了污染数据集的构建方法。代码与数据可在https://github.com/HoyaAm/Double-Landmines 获取。

原文摘要 · Abstract (English)

Backdoor attacks pose an important security threat to textual large language models. Exploring textual backdoor attacks not only helps reveal the potential security risks of models, but also promotes innovation and development of defense mechanisms. Currently, most textual backdoor attack methods are based on a single trigger. For example, inserting specific content into text as a trigger or changing the abstract text features to be a trigger. However, the adoption of this single-trigger mode makes the existing backdoor attacks subject to certain limitations: either they are easily identified by the existing defense strategies, or they have certain shortcomings in attack performance and in the construction of poisoned datasets. In order to solve these issues, a dual-trigger backdoor attack method is proposed in this paper. Specifically, we use two different attributes, syntax and mood (we use subjunctive mood as an example in this article), as two different triggers. It makes our backdoor attack method similar to a double landmine which can have completely different trigger conditions simultaneously. Therefore, this method not only improves the flexibility of trigger mode, but also enhances the robustness against defense detection. A large number of experimental results show that this method significantly outperforms the previous methods based on abstract features in attack performance, and achieves comparable attack performance (almost 100\% attack success rate) with the insertion-based method. In addition, in order to further improve the attack performance, we also give the construction method of the poisoned dataset.The code and data of this paper can be obtained at https://github.com/HoyaAm/Double-Landmines.

文本安全后门攻击大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。