arXiv:2412.12192cs.CRcs.AI2024-12被引 1

用特定句式示范可有效防御预填充攻击,但会过度防御。

No Free Lunch for Defending Against Prefilling Attack by In-Context Learning

  • 在示例中加入对抗性句式,利用ICL实现防御。
  • 不同模型规模下均有效,但存在过度防御现象。
  • 适合关注LLM安全与对抗样本的研究者。

大型语言模型(LLMs)的安全性自ChatGPT问世以来成为重要研究方向。尽管已有多种有效方法防御越狱攻击,预填充攻击仍是开源LLM面临的重要且未解决的威胁。上下文学习(In-Context Learning, ICL)能高效防御多种越狱攻击,但尚无有效ICL方法应对预填充攻击。本文表明:(1) 通过在演示中使用对抗性句式结构,ICL可有效防御预填充越狱攻击;(2) 从模型规模、演示数量、过度防御、与其他越狱攻击的结合、安全对齐的存在性等角度分析该防御的有效性。实验结果表明,使用ICL防御预填充攻击不存在“免费午餐”:当前安全对齐方法无法缓解此类攻击,而对抗性结构的ICL示范在各类模型规模和复杂越狱攻击下表现稳健;然而,模型在使用此类示范时表现出相似的过度防御行为,且该行为与模型规模无关。

原文摘要 · Abstract (English)

The security of Large Language Models (LLMs) has become an important research topic since the emergence of ChatGPT. Though there have been various effective methods to defend against jailbreak attacks, prefilling attacks remain an unsolved and popular threat against open-sourced LLMs. In-Context Learning (ICL) offers a computationally efficient defense against various jailbreak attacks, yet no effective ICL methods have been developed to counter prefilling attacks. In this paper, we: (1) show that ICL can effectively defend against prefilling jailbreak attacks by employing adversative sentence structures within demonstrations; (2) characterize the effectiveness of this defense through the lens of model size, number of demonstrations, over-defense, integration with other jailbreak attacks, and the presence of safety alignment. Given the experimental results and our analysis, we conclude that there is no free lunch for defending against prefilling jailbreak attacks with ICL. On the one hand, current safety alignment methods fail to mitigate prefilling jailbreak attacks, but adversative structures within ICL demonstrations provide robust defense across various model sizes and complex jailbreak attacks. On the other hand, LLMs exhibit similar over-defensiveness when utilizing ICL demonstrations with adversative structures, and this behavior appears to be independent of model size.

大模型安全对抗攻击上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。