用自适应优化机制生成更隐蔽的后门文本,提升攻击效果与质量。
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
- 通过迭代优化提示词,让大模型自动生成高质量毒化文本。
- 在六种攻击、两种防御下平均攻击成功率仍达96.75%。
- 无需人工设计提示,适配性强,适合研究后门防御的学者。
现有基于插入和改写方式的后门攻击虽有效,但忽视了毒化文本与正常文本在语义一致性和质量上的差异。尽管近期研究利用大模型生成毒化文本以提升隐蔽性与质量,其手工设计的提示依赖专家经验,存在提示泛化能力差、防御后性能下降的问题。本文提出一种基于黑盒大模型自适应优化机制的新型后门攻击方法 BadApex,利用黑盒大模型通过优化提示词生成毒化文本。具体地,设计自适应优化机制,由生成代理基于初始提示生成毒化文本,修改代理评估文本质量并优化新提示,经多轮迭代后使用最终提示生成毒化文本。在三个数据集上开展六种攻击与两种防御的实验,结果表明,BadApex显著优于现有最先进攻击方法,在提示适应性、语义一致性与文本质量方面均有提升;即使面对双重防御,平均攻击成功率仍高达96.75%。
原文摘要 · Abstract (English)
Previous insertion-based and paraphrase-based backdoors have achieved great success in attack efficacy, but they ignore the text quality and semantic consistency between poisoned and clean texts. Although recent studies introduce LLMs to generate poisoned texts and improve the stealthiness, semantic consistency, and text quality, their hand-crafted prompts rely on expert experiences, facing significant challenges in prompt adaptability and attack performance after defenses. In this paper, we propose a novel backdoor attack based on adaptive optimization mechanism of black-box large language models (BadApex), which leverages a black-box LLM to generate poisoned text through a refined prompt. Specifically, an Adaptive Optimization Mechanism is designed to refine an initial prompt iteratively using the generation and modification agents. The generation agent generates the poisoned text based on the initial prompt. Then the modification agent evaluates the quality of the poisoned text and refines a new prompt. After several iterations of the above process, the refined prompt is used to generate poisoned texts through LLMs. We conduct extensive experiments on three dataset with six backdoor attacks and two defenses. Extensive experimental results demonstrate that BadApex significantly outperforms state-of-the-art attacks. It improves prompt adaptability, semantic consistency, and text quality. Furthermore, when two defense methods are applied, the average attack success rate (ASR) still up to 96.75%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。