arXiv:2507.02956cs.CRcs.AI2025-07被引 10

揭示多轮越狱攻击如何欺骗大模型的内部表征,使其误判有害请求为安全内容。

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks

  • 从模型中间表示层面分析多轮越狱攻击机制
  • 越狱请求在多轮对话中持续保持在安全表征区域
  • 解释单轮防御失效原因,适合安全研究者参考

近期研究显示,当前最先进大模型及其防御机制仍易受多轮越狱攻击。此类攻击仅需封闭式模型访问权限,且常可手动完成,对基于大模型系统的安全部署构成重大威胁。本文从中间模型表征层面研究Crescendo多轮越狱攻击的有效性,发现安全对齐的大模型往往将Crescendo生成的响应视为较无害,尤其随着对话轮数增加而愈发明显。分析表明,每一轮中,Crescendo提示均使模型输出维持在表示空间中的‘良性’区域,从而有效诱骗模型执行有害请求。此外,研究结果有助于解释为何单轮越狱防御(如电路断路器)普遍无法应对多轮攻击,推动针对此泛化差距的缓解策略发展。

原文摘要 · Abstract (English)

Recent research has demonstrated that state-of-the-art LLMs and defenses remain susceptible to multi-turn jailbreak attacks. These attacks require only closed-box model access and are often easy to perform manually, posing a significant threat to the safe and secure deployment of LLM-based systems. We study the effectiveness of the Crescendo multi-turn jailbreak at the level of intermediate model representations and find that safety-aligned LMs often represent Crescendo responses as more benign than harmful, especially as the number of conversation turns increases. Our analysis indicates that at each turn, Crescendo prompts tend to keep model outputs in a "benign" region of representation space, effectively tricking the model into fulfilling harmful requests. Further, our results help explain why single-turn jailbreak defenses like circuit breakers are generally ineffective against multi-turn attacks, motivating the development of mitigations that address this generalization gap.

大模型安全越狱攻击表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。