arXiv:2606.00150cs.CRcs.AI2026-06

通过逐步注入指令,让大模型遗忘安全机制,成功率高达95%。

Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models

论文配图:Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
图 1 · 摘自论文原文
  • 分步注入指令,利用模型记忆特性突破安全限制。
  • 累积注入后,攻击成功率最高达95%。
  • 揭示模型记忆机制漏洞,适合安全研究者关注。

随着大语言模型为用户便利不断演进,尽管持续进行安全训练,其仍易受越狱攻击。传统越狱方法通常仅针对单一提示注入,忽视了模型对对话流程和用户指令的记忆能力。本文提出Persona Attack,一种基于记忆注入的越狱方法,通过逐步操作影响模型的上下文窗口。实验表明,随着注入指令在记忆中积累,模型逐渐优先执行这些指令,而非内部的安全对齐机制。此外,实证结果显示,攻击成功率不仅与模型的记忆实现方式有关,还受指令组合影响,在特定配置下可达95%。

原文摘要 · Abstract (English)

As Large Language Models evolve for user convenience, vulnerability to jailbreak attacks continues to be reported despite ongoing efforts in safety training. Traditional jailbreak techniques typically focus on a single prompt injection, neglecting the models' ability to remember the flow of conversation and the user's instructions. In this paper, we propose Persona Attack, a memory injection based jailbreak method that manipulates the model's context window through a step by step approach. Experimental results from applying Persona Attack to several widely used LLMs reveal that, as injections accumulate in memory, models increasingly prioritize these instructions over their internal safety alignment mechanisms. Furthermore, our experiments empirically demonstrate that the attack success rate varies not only according to the memory implementation of the model, but also combinations of instructions and can reach 95% under specific instruction configurations.

越狱攻击记忆注入LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。