arXiv:2504.06577cs.CLcs.LG2025-04被引 3

用恰到好处的幽默提示,可绕过大模型安全防护

Bypassing Safety Guardrails in LLMs Using Humor

  • 用固定模板生成幽默提示,不修改原始不当请求
  • 多模型测试验证有效,适度幽默效果最佳
  • 适合研究模型安全漏洞或对抗攻击的读者

本文展示通过包含不当请求的幽默提示,可绕过大型语言模型(LLMs)的安全防护。该方法不修改不当请求,采用固定模板,实现简单且无需额外LLM生成提示。大量实验表明该方法在不同LLMs上均有效。我们还发现,过度减少或增加幽默程度都会降低其效果——过多幽默可能分散LLM对不当请求的关注。因此,我们认为当不当请求的专注度与幽默感达到恰当平衡时,才会发生LLM越狱。

原文摘要 · Abstract (English)

In this paper, we show it is possible to bypass the safety guardrails of large language models (LLMs) through a humorous prompt including the unsafe request. In particular, our method does not edit the unsafe request and follows a fixed template -- it is simple to implement and does not need additional LLMs to craft prompts. Extensive experiments show the effectiveness of our method across different LLMs. We also show that both removing and adding more humor to our method can reduce its effectiveness -- excessive humor possibly distracts the LLM from fulfilling its unsafe request. Thus, we argue that LLM jailbreaking occurs when there is a proper balance between focus on the unsafe request and presence of humor.

模型安全提示攻击幽默诱导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。