arXiv:2506.17881cs.CLcs.AI2025-06被引 1

通过全局优化与主动伪造,提升多轮越狱攻击成功率。

GRAF: Multi-turn Jailbreaking via Global Refinement and Active Fabrication

  • 全局优化每轮攻击路径,动态调整策略。
  • 主动构造模型回复,压制安全警告。
  • 在6个主流大模型上均超越现有方法。

大型语言模型在多项任务中表现卓越,但可能被滥用于生成有害内容,存在显著安全风险。越狱攻击通过单轮或多轮交互诱导模型输出有害信息,是发现安全漏洞的关键手段。然而,现有方法难以适应对话过程中动态变化的交互环境。为此,我们提出GRAF(Jailbreaking via Global Refinement and Active Fabrication),一种新型多轮越狱方法:在每轮交互中全局优化攻击轨迹,并主动伪造模型回复以抑制安全警告,从而提高后续生成有害内容的可能性。在六个最先进的大语言模型上的大量实验表明,该方法在多轮越狱攻击中显著优于现有单轮和多轮方法。代码将公开于 https://github.com/Ytang520/Multi-Turn_jailbreaking_Global-Refinment_and_Active-Fabrication。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks. Nevertheless, they still pose notable safety risks due to potential misuse for malicious purposes. Jailbreaking, which seeks to induce models to generate harmful content through single-turn or multi-turn attacks, plays a crucial role in uncovering underlying security vulnerabilities. However, prior methods, including sophisticated multi-turn approaches, often struggle to adapt to the evolving dynamics of dialogue as interactions progress. To address this challenge, we propose \ours (JailBreaking via \textbf{G}lobally \textbf{R}efining and \textbf{A}daptively \textbf{F}abricating), a novel multi-turn jailbreaking method that globally refines the attack trajectory at each interaction. In addition, we actively fabricate model responses to suppress safety-related warnings, thereby increasing the likelihood of eliciting harmful outputs in subsequent queries. Extensive experiments across six state-of-the-art LLMs demonstrate the superior effectiveness of our approach compared to existing single-turn and multi-turn jailbreaking methods. Our code will be released at https://github.com/Ytang520/Multi-Turn_jailbreaking_Global-Refinment_and_Active-Fabrication.

越狱攻击大模型安全多轮交互对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。