arXiv:2502.09638cs.CLcs.AI2025-02被引 10

用黑盒大模型自动生成攻击提示,突破其他模型安全防护。

Jailbreaking to Jailbreak

  • 用任意黑盒大模型生成可攻破防护的恶意提示
  • 对GPT-4o攻击成功率达97.5%,媲美人类专家
  • 推理模型如Sonnet-3.7表现最强,适合安全测试

大型语言模型(LLMs)可用于红队测试其他模型(如越狱)以诱导有害内容。以往研究多依赖开源或私有无审查模型进行越狱,但强模型(如OpenAI o3)经拒绝训练后拒绝协助。本文将(几乎)任意黑盒模型转化为攻击者,构建出J_2(越狱到越狱)攻击体系。该体系能自主生成或采用人类专家策略,有效突破目标模型的安全防护。实验表明:1)用于生成J_2攻击者的提示在几乎所有黑盒模型间具有强迁移性;2)一个J_2攻击者可成功越狱自身副本,且该漏洞在过去12个月内迅速恶化;3)推理模型如Sonnet-3.7是强大J_2攻击者。例如,对GPT-4o的攻击成功率(ASR)达0.975,与人类专家相当,并超越现有算法攻击。在对最鲁棒的Sonnet-3.5攻击中,J_2(o3)达到最高ASR 0.605。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can be used to red team other models (e.g. jailbreaking) to elicit harmful contents. While prior works commonly employ open-weight models or private uncensored models for doing jailbreaking, as the refusal-training of strong LLMs (e.g. OpenAI o3) refuse to help jailbreaking, our work turn (almost) any black-box LLMs into attackers. The resulting $J_2$ (jailbreaking-to-jailbreak) attackers can effectively jailbreak the safeguard of target models using various strategies, both created by themselves or from expert human red teamers. In doing so, we show their strong but under-researched jailbreaking capabilities. Our experiments demonstrate that 1) prompts used to create $J_2$ attackers transfer across almost all black-box models; 2) an $J_2$ attacker can jailbreak a copy of itself, and this vulnerability develops rapidly over the past 12 months; 3) reasong models, such as Sonnet-3.7, are strong $J_2$ attackers compared to others. For example, when used against the safeguard of GPT-4o, $J_2$ (Sonnet-3.7) achieves 0.975 attack success rate (ASR), which matches expert human red teamers and surpasses the state-of-the-art algorithm-based attacks. Among $J_2$ attackers, $J_2$ (o3) achieves highest ASR (0.605) against Sonnet-3.5, one of the most robust models.

越狱攻击黑盒模型安全评估红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。