arXiv:2605.01687cs.CL2026-05被引 5

构建首个大规模多轮越狱攻击基准,揭示真实对话中模型安全漏洞。

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

论文配图:MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
图 1 · 摘自论文原文
  • 基于主动学习生成多样高质多轮越狱提示,持续优化攻击强度。
  • 含10,389条多轮攻击样本,覆盖2,665种有害意图,成功率超现有数据集54%。
  • 发现单轮看似无害的攻击在多轮中威胁加剧,适合安全研究者使用。

我们提出 MultiBreak,一个可扩展且多样化的多轮越狱攻击评估基准,用于评测大语言模型(LLM)的安全性。相比单轮攻击,多轮越狱更贴近自然对话场景,更易绕过安全对齐的 LLM。现有多轮基准规模有限或高度依赖模板,缺乏多样性。为弥补这一差距,我们整合多种有害越狱意图,并引入一种主动学习流水线,通过不确定性驱动的精炼策略,迭代微调生成器以产出更强攻击候选。MultiBreak 包含 10,389 条多轮对抗性提示,涵盖 2,665 种不同有害意图,是迄今主题最广泛的集合。实证评估显示,该基准在 DeepSeek-R1-7B 和 GPT-4.1-mini 上的攻击成功率(ASR)分别比第二优数据集高出 54.0% 和 34.6%。更重要的是,安全评估表明,多样化攻击类别能暴露模型的细粒度漏洞;某些在单轮下看似无害的攻击,在多轮场景中表现出显著更高的攻击效力。这些发现凸显了模型在真实对抗环境下的持续脆弱性,并确立 MultiBreak 作为推动 LLM 安全研究的可扩展资源。

原文摘要 · Abstract (English)

We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic natural conversational settings, making them easier to bypass safety-aligned LLM than single-turn jailbreaks. Existing multi-turn benchmarks are limited in size or rely heavily on templates, which restrict their diversity. To address this gap, we unify a wide range of harmful jailbreak intents, and introduce an active learning pipeline for expanding high-quality multi-turn adversarial prompts, where a generator is iteratively fine-tuned to produce stronger attack candidates, guided by uncertainty-based refinement. Our MultiBreak includes 10,389 multi-turn adversarial prompts, spans 2,665 distinct harmful intents, and covers the most diverse set of topics to date. Empirical evaluation shows that our benchmark achieves up to a 54.0 and 34.6 higher attack success rate (ASR)} than the second-best dataset on DeepSeek-R1-7B and GPT-4.1-mini, respectively. More importantly, safety evaluations suggest that diverse attack categories uncover fine-grained LLM vulnerabilities}, and categories that appear benign under single-turn can exhibit substantially higher adversarial effectiveness in multi-turn scenarios. These findings highlight persistent vulnerabilities of LLMs under realistic adversarial settings and establish MultiBreak as a scalable resource for advancing LLM safety.

LLM安全越狱攻击多轮对话基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。