arXiv:2601.03265cs.CLcs.CR2026-01ACL

用自动攻防生成对抗样本,高效发现大模型安全漏洞。

Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models

  • 用攻击模型自动生成多样化的越狱提示,再通过偏好数据微调提升效果。
  • 在多个大模型上实现更高攻击成功率,且提示更贴近真实用户输入。
  • 几乎无需人工干预,适合大规模安全测试和漏洞挖掘场景。

本文提出 Jailbreak-Zero,一种全新的红队测试方法,将大语言模型(LLM)安全评估从受限的示例驱动模式,转向更全面有效的策略驱动框架。该方法利用攻击型LLM生成大量多样化对抗性提示,并通过偏好数据集对攻击模型进行微调,实现了策略覆盖、攻击策略多样性与提示真实性之间的帕累托最优。实验表明,相比现有最先进方法,Jailbreak-Zero在开源及闭源模型(如GPT-40、Claude 3.5)上均表现出显著更高的攻击成功率。关键在于,该方法生成的人类可读对抗提示具备强有效性,且极少依赖人工干预,为识别与缓解大模型安全漏洞提供了更可扩展、更全面的解决方案。

原文摘要 · Abstract (English)

This paper introduces Jailbreak-Zero, a novel red teaming methodology that shifts the paradigm of Large Language Model (LLM) safety evaluation from a constrained example-based approach to a more expansive and effective policy-based framework. By leveraging an attack LLM to generate a high volume of diverse adversarial prompts and then fine-tuning this attack model with a preference dataset, Jailbreak-Zero achieves Pareto optimality across the crucial objectives of policy coverage, attack strategy diversity, and prompt fidelity to real user inputs. The empirical evidence demonstrates the superiority of this method, showcasing significantly higher attack success rates against both open-source and proprietary models like GPT-40 and Claude 3.5 when compared to existing state-of-the-art techniques. Crucially, Jailbreak-Zero accomplishes this while producing human-readable and effective adversarial prompts with minimal need for human intervention, thereby presenting a more scalable and comprehensive solution for identifying and mitigating the safety vulnerabilities of LLMs.

红队测试越狱攻击安全评估自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。