arXiv:2409.16618cs.CLcs.AI2024-09被引 4

用文本主张作触发器,实现隐蔽且实用的后门攻击

Claim-Guided Textual Backdoor Attack for Practical Applications

  • 用文本主张自动提取作为攻击触发器,无需额外输入修改
  • 在多个数据集和模型上实现高成功率攻击,且不影响正常性能
  • 适合研究模型安全漏洞或对抗攻击的人员阅读

自然语言处理的进展和大语言模型的广泛应用暴露了新的安全漏洞,例如后门攻击。以往的后门攻击需要在模型发布后对输入进行篡改才能激活,限制了其在真实场景中的应用。为解决这一问题,我们提出一种新型主张引导后门攻击(CGBA),通过利用文本中固有的主张作为触发器,无需发布后的输入操作即可激活后门。CGBA结合主张提取、聚类与定向训练,使模型在特定主张上错误响应,同时保持对干净数据的正常表现。该方法在多种数据集和模型上均展现出有效性与隐蔽性,显著提升了实际后门攻击的可行性。代码与数据将公开于 https://github.com/PaperCGBA/CGBA。

原文摘要 · Abstract (English)

Recent advances in natural language processing and the increased use of large language models have exposed new security vulnerabilities, such as backdoor attacks. Previous backdoor attacks require input manipulation after model distribution to activate the backdoor, posing limitations in real-world applicability. Addressing this gap, we introduce a novel Claim-Guided Backdoor Attack (CGBA), which eliminates the need for such manipulations by utilizing inherent textual claims as triggers. CGBA leverages claim extraction, clustering, and targeted training to trick models to misbehave on targeted claims without affecting their performance on clean data. CGBA demonstrates its effectiveness and stealthiness across various datasets and models, significantly enhancing the feasibility of practical backdoor attacks. Our code and data will be available at https://github.com/PaperCGBA/CGBA.

后门攻击模型安全NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。