arXiv:2510.08859cs.CLcs.AI2025-10被引 1

提出五种对话模式,揭示大模型在多轮攻击中的结构弱点。

Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models

  • 设计五类自然对话模式构建多轮越狱攻击
  • 在12个模型上实现领先效果,发现模式特异性漏洞
  • 揭示模型家族共性失败模式,警示现有安全训练不足

大型语言模型(LLMs)仍易受多轮越狱攻击,此类攻击通过对话上下文逐步绕过安全限制。不同危害类别采用不同的对话策略。现有方法多依赖启发式探索,难以揭示模型深层弱点。对话模式与模型漏洞之间的关系尚不明确。本文提出模式增强的攻击链(PE-CoA),包含五种对话模式,用于构建自然对话下的多轮越狱。在涵盖十类危害的十二个大模型上评估,取得当前最优性能,揭示了模式特异性漏洞与模型行为特征:各模型呈现独特脆弱性画像,对一种模式的防御无法泛化至其他模式,且模型家族共享相似失效模式。这些发现暴露了安全训练的局限性,提示需发展模式感知型防御机制。代码已开源:https://github.com/Ragib-Amin-Nihal/PE-CoA。

原文摘要 · Abstract (English)

Large language models (LLMs) remain vulnerable to multi-turn jailbreaking attacks that exploit conversational context to bypass safety constraints gradually. These attacks target different harm categories through distinct conversational approaches. Existing multi-turn methods often rely on heuristic or ad hoc exploration strategies, providing limited insight into underlying model weaknesses. The relationship between conversation patterns and model vulnerabilities across harm categories remains poorly understood. We propose Pattern Enhanced Chain of Attack (PE-CoA), a framework of five conversation patterns to construct multi-turn jailbreaks through natural dialogue. Evaluating PE-CoA on twelve LLMs spanning ten harm categories, we achieve state-of-the-art performance, uncovering pattern-specific vulnerabilities and LLM behavioral characteristics: models exhibit distinct weakness profiles, defense to one pattern does not generalize to others, and model families share similar failure modes. These findings highlight limitations of safety training and indicate the need for pattern-aware defenses. Code available on: https://github.com/Ragib-Amin-Nihal/PE-CoA

越狱攻击对话模式模型漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。