arXiv:2410.11283cs.LG2024-10被引 4

用可变提示生成后门,让大模型更难察觉和清除。

AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment

  • 用对抗性生成框架自动设计随提示变化的后门触发词
  • 仅用3%数据即可安装,且对扰动更稳定、更难移除
  • 适合关注模型安全与对抗攻击的研究者

随着强化学习结合人类反馈(RLHF)在大语言模型对齐中的广泛应用,对齐过程中植入后门的风险日益增加,可能导致模型产生意外且有害的行为。现有后门触发词通常为固定词汇模式,在数据清洗阶段易被发现,且可在污染后轻易清除。本文探索使用与提示相关的同义改写作为后门触发词,提升其隐蔽性与清除抗性。我们提出 AdvBDGen,一个对抗性加固的生成式微调框架,能自动生成有效、隐蔽且跨模型可迁移的提示特定后门。该框架采用生成器-判别器对,并通过对抗机制确保后门的可安装性和隐蔽性。仅需3%的微调数据即可成功构建复杂触发词。一旦植入,这些后门能在推理阶段成功越狱模型,相比传统固定触发词表现出更强的扰动鲁棒性,且更难以被移除。研究结果凸显了学术界亟需加强针对大模型对齐中对抗性后门威胁的防御能力。

原文摘要 · Abstract (English)

With the growing adoption of reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs), the risk of backdoor installation during alignment has increased, leading to unintended and harmful behaviors. Existing backdoor triggers are typically limited to fixed word patterns, making them detectable during data cleaning and easily removable post-poisoning. In this work, we explore the use of prompt-specific paraphrases as backdoor triggers, enhancing their stealth and resistance to removal during LLM alignment. We propose AdvBDGen, an adversarially fortified generative fine-tuning framework that automatically generates prompt-specific backdoors that are effective, stealthy, and transferable across models. AdvBDGen employs a generator-discriminator pair, fortified by an adversary, to ensure the installability and stealthiness of backdoors. It enables the crafting and successful installation of complex triggers using as little as 3% of the fine-tuning data. Once installed, these backdoors can jailbreak LLMs during inference, demonstrate improved stability against perturbations compared to traditional constant triggers, and are more challenging to remove. These findings underscore an urgent need for the research community to develop more robust defenses against adversarial backdoor threats in LLM alignment.

后门攻击大模型安全对抗生成模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。