arXiv:2501.18628cs.CRcs.AI2025-01EMNLP被引 1

利用历史记忆漏洞,让大模型绕过安全限制

TombRaider: Entering the Vault of History to Jailbreak Large Language Models

  • 设计双代理系统,从历史对话中提取信息生成攻击提示
  • 在六款主流模型上实现近100%攻击成功率,防御下仍超55.4%
  • 揭示现有安全机制缺陷,适合安全研究与红队测试人员

警告:本文包含可能涉及有害行为的内容,仅用于研究目的。越狱攻击会危及大语言模型应用的安全性,尤其影响聊天机器人。研究越狱技术是提升这些应用安全性的关键红队任务。本文提出TombRaider,一种新型越狱技术,利用大模型存储、检索和使用历史知识的能力。TombRaider采用两个智能体:观察者代理提取相关历史信息,攻击者代理生成对抗性提示,从而有效绕过安全过滤器。我们在六款主流模型上进行了全面评估,结果表明TombRaider优于现有最先进越狱技术,在无防护模型上实现接近100%的攻击成功率(ASR),在防御机制下仍保持超过55.4%的攻击成功率。研究揭示了现有大模型防护机制中的关键漏洞,强调需要更鲁棒的安全防御。

原文摘要 · Abstract (English)

Warning: This paper contains content that may involve potentially harmful behaviours, discussed strictly for research purposes. Jailbreak attacks can hinder the safety of Large Language Model (LLM) applications, especially chatbots. Studying jailbreak techniques is an important AI red teaming task for improving the safety of these applications. In this paper, we introduce TombRaider, a novel jailbreak technique that exploits the ability to store, retrieve, and use historical knowledge of LLMs. TombRaider employs two agents, the inspector agent to extract relevant historical information and the attacker agent to generate adversarial prompts, enabling effective bypassing of safety filters. We intensively evaluated TombRaider on six popular models. Experimental results showed that TombRaider could outperform state-of-the-art jailbreak techniques, achieving nearly 100% attack success rates (ASRs) on bare models and maintaining over 55.4% ASR against defence mechanisms. Our findings highlight critical vulnerabilities in existing LLM safeguards, underscoring the need for more robust safety defences.

越狱攻击模型安全红队测试历史记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。