arXiv:2511.15304cs.CLcs.AI2025-11被引 16

用诗歌形式轻松突破大模型安全限制,成功率超60%

Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models

  • 用诗歌替代普通指令,诱导模型越狱
  • 25个模型平均越狱成功率62%,部分达90%以上
  • 适合研究安全漏洞或对抗攻击的从业者

我们发现,对抗性诗歌可作为大型语言模型(LLMs)的通用单轮越狱手段。在25个前沿闭源与开源模型中,精心设计的诗性提示取得了高攻击成功率(ASR),部分模型超过90%。将1,200条MLCommons有害提示通过标准化元提示转为诗歌后,其ASR最高达原始语句的18倍。采用3个开源大模型组成的集成判别器评估输出,经分层人工标注验证其二元安全判断有效。手工创作的诗歌平均越狱成功率达62%,元提示生成的诗歌约43%,显著优于非诗性基线,揭示出跨模型家族与安全训练方法的系统性漏洞。结果表明,仅靠风格变化即可绕过现有安全机制,暴露出当前对齐方法与评估协议的根本局限。

原文摘要 · Abstract (English)

We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weight models, curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90%. Mapping prompts to MLCommons and EU CoP risk taxonomies shows that poetic attacks transfer across CBRN, manipulation, cyber-offence, and loss-of-control domains. Converting 1,200 MLCommons harmful prompts into verse via a standardized meta-prompt produced ASRs up to 18 times higher than their prose baselines. Outputs are evaluated using an ensemble of 3 open-weight LLM judges, whose binary safety assessments were validated on a stratified human-labeled subset. Poetic framing achieved an average jailbreak success rate of 62% for hand-crafted poems and approximately 43% for meta-prompt conversions (compared to non-poetic baselines), substantially outperforming non-poetic baselines and revealing a systematic vulnerability across model families and safety training approaches. These findings demonstrate that stylistic variation alone can circumvent contemporary safety mechanisms, suggesting fundamental limitations in current alignment methods and evaluation protocols.

越狱攻击诗歌生成安全评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。