arXiv:2601.05742cs.CRcs.AI2026-01被引 4

提出新型多轮越狱攻击Echo Chamber,逐步突破聊天机器人安全防护。

The Echo Chamber Multi-Turn LLM Jailbreak

  • 通过渐进式交互链实现多轮越狱,绕过安全限制。
  • 在多个顶尖模型上验证有效,成功率显著高于现有方法。
  • 适合研究模型安全与防御机制的人员参考。

大型语言模型(LLMs)的普及催生了低成本开发的强大聊天机器人。随着企业部署这些工具,需应对安全挑战以避免财务损失和声誉损害。其中关键问题是越狱攻击,即恶意操控提示词和输入以绕过聊天机器人的安全防护机制。多轮攻击是一种新型越狱方式,涉及与聊天机器人精心设计的连续交互链。本文提出新的多轮攻击方法Echo Chamber,采用渐进式升级策略,详细描述其原理,并与其它多轮攻击进行对比。通过大规模评估,证明该方法在多个最先进的模型上均表现出色,有效突破安全防线。

原文摘要 · Abstract (English)

The availability of Large Language Models (LLMs) has led to a new generation of powerful chatbots that can be developed at relatively low cost. As companies deploy these tools, security challenges need to be addressed to prevent financial loss and reputational damage. A key security challenge is jailbreaking, the malicious manipulation of prompts and inputs to bypass a chatbot's safety guardrails. Multi-turn attacks are a relatively new form of jailbreaking involving a carefully crafted chain of interactions with a chatbot. We introduce Echo Chamber, a new multi-turn attack using a gradual escalation method. We describe this attack in detail, compare it to other multi-turn attacks, and demonstrate its performance against multiple state-of-the-art models through extensive evaluation.

越狱攻击多轮攻击安全防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。