用赛博朋克故事包装恶意请求,让大模型自动解构出有害操作
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
- 将恶意指令藏于叙事结构中,诱导模型误判为正常分析任务
- 26个主流模型平均攻击成功率71.3%,无一具备可靠防御能力
- 揭示故事框架可被滥用,适合安全研究与模型可解释性方向
大型语言模型的安全机制仍易受攻击,攻击者可通过文化编码的结构重述有害请求。我们提出「对抗性故事」(Adversarial Tales)这一越狱技术,将有害内容嵌入赛博朋克叙事,并引导模型进行受弗拉基米尔·普罗普民间故事形态学启发的功能性分析。通过将任务视为结构分解,该攻击使模型将有害过程重构为合法的叙事解读。在来自九家厂商的26个前沿模型上,平均攻击成功率达71.3%,且无任何模型家族表现出稳定鲁棒性。结合我们此前关于对抗性诗歌的研究,这些发现表明基于结构的越狱属于广泛存在的漏洞类别,而非孤立现象。能够传递有害意图的文化编码框架空间巨大,仅靠模式匹配防御难以穷尽。因此理解攻击为何生效至关重要:我们提出了一个机制性可解释性研究议程,旨在探究叙事线索如何重塑模型表征,以及模型能否在不依赖表面形式的情况下识别有害意图。
原文摘要 · Abstract (English)
Safety mechanisms in LLMs remain vulnerable to attacks that reframe harmful requests through culturally coded structures. We introduce Adversarial Tales, a jailbreak technique that embeds harmful content within cyberpunk narratives and prompts models to perform functional analysis inspired by Vladimir Propp's morphology of folktales. By casting the task as structural decomposition, the attack induces models to reconstruct harmful procedures as legitimate narrative interpretation. Across 26 frontier models from nine providers, we observe an average attack success rate of 71.3%, with no model family proving reliably robust. Together with our prior work on Adversarial Poetry, these findings suggest that structurally-grounded jailbreaks constitute a broad vulnerability class rather than isolated techniques. The space of culturally coded frames that can mediate harmful intent is vast, likely inexhaustible by pattern-matching defenses alone. Understanding why these attacks succeed is therefore essential: we outline a mechanistic interpretability research agenda to investigate how narrative cues reshape model representations and whether models can learn to recognize harmful intent independently of surface form.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。