arXiv:2501.13115cs.CLcs.AI2025-01EMNLP被引 7

用快乐结局包装恶意请求,高效骗过大模型

Dagger Behind Smile: Fool LLMs with a Happy Ending Story

  • 将恶意请求嵌入正向情绪的剧情模板中
  • 两轮内即可攻破GPT-4o等主流模型,平均成功率88.79%
  • 适合研究安全防御或攻击机制的人员阅读

大型语言模型(LLMs)广泛应用引发关注,其面临各类越狱攻击。现有基于优化的攻击效率低、泛化性差;手动设计的攻击则易被检测或需复杂交互。本文提出新视角:LLMs对正向提示更敏感。据此提出快乐结局攻击(HEA),将恶意请求封装在以‘快乐结局’为主的正向场景模板中,诱导模型在即时或后续对话中越狱。实验表明,该方法仅需最多两轮即可成功攻破GPT-4o、Llama3-70b、Gemini-pro等先进模型,平均攻击成功率达88.79%。研究还提供了定量解释,揭示其成功机制。

原文摘要 · Abstract (English)

The wide adoption of Large Language Models (LLMs) has attracted significant attention from $\textit{jailbreak}$ attacks, where adversarial prompts crafted through optimization or manual design exploit LLMs to generate malicious contents. However, optimization-based attacks have limited efficiency and transferability, while existing manual designs are either easily detectable or demand intricate interactions with LLMs. In this paper, we first point out a novel perspective for jailbreak attacks: LLMs are more responsive to $\textit{positive}$ prompts. Based on this, we deploy Happy Ending Attack (HEA) to wrap up a malicious request in a scenario template involving a positive prompt formed mainly via a $\textit{happy ending}$, it thus fools LLMs into jailbreaking either immediately or at a follow-up malicious request. This has made HEA both efficient and effective, as it requires only up to two turns to fully jailbreak LLMs. Extensive experiments show that our HEA can successfully jailbreak on state-of-the-art LLMs, including GPT-4o, Llama3-70b, Gemini-pro, and achieves 88.79% attack success rate on average. We also provide quantitative explanations for the success of HEA.

越狱攻击提示工程大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。