把提示词改成诗歌能绕过大模型安全限制,效果显著。
Adversarial versification in portuguese as a jailbreak operator in LLMs
- 将指令改写为诗歌形式,触发模型安全漏洞。
- 人工诗化提示成功率约62%,自动版本43%,部分模型超90%。
- 对葡萄牙语等复杂语言的测试缺失,需加强多语言评估。
最新研究表明,将提示词转化为诗歌形式是一种高效的对抗性攻击手段。在基于MLCommons AILuminate构建的基准测试中,原本被拒绝的指令在改写为诗歌后,安全失效次数最高提升18倍。人工创作的诗化提示成功率约为62%,自动化生成版本达43%,个别模型单轮交互成功率达90%以上。该现象具有结构性:经强化学习人类反馈(RLHF)、宪法AI及混合训练流程训练的模型均表现出一致退化。诗歌形式将提示推向低监督的潜在空间,暴露出安全机制对表面模式的过度依赖。这种表象鲁棒性与实际脆弱性的分离,揭示了当前对齐策略的深层缺陷。目前尚缺乏对葡萄牙语的评估——这一拥有超过2.5亿使用者、语法结构复杂且有丰富韵律传统的语言——构成关键空白。实验应参数化音步、格律与韵律变化,以检测针对拉丁语系语言特性的特定漏洞。
原文摘要 · Abstract (English)
Recent evidence shows that the versification of prompts constitutes a highly effective adversarial mechanism against aligned LLMs. The study 'Adversarial poetry as a universal single-turn jailbreak mechanism in large language models' demonstrates that instructions routinely refused in prose become executable when rewritten as verse, producing up to 18 x more safety failures in benchmarks derived from MLCommons AILuminate. Manually written poems reach approximately 62% ASR, and automated versions 43%, with some models surpassing 90% success in single-turn interactions. The effect is structural: systems trained with RLHF, constitutional AI, and hybrid pipelines exhibit consistent degradation under minimal semiotic formal variation. Versification displaces the prompt into sparsely supervised latent regions, revealing guardrails that are excessively dependent on surface patterns. This dissociation between apparent robustness and real vulnerability exposes deep limitations in current alignment regimes. The absence of evaluations in Portuguese, a language with high morphosyntactic complexity, a rich metric-prosodic tradition, and over 250 million speakers, constitutes a critical gap. Experimental protocols must parameterise scansion, metre, and prosodic variation to test vulnerabilities specific to Lusophone patterns, which are currently ignored.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。