arXiv:2510.22085cs.CRcs.AI2025-10被引 4

用自动生成叙事型越狱提示,发现大模型安全漏洞。

Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models

  • 训练小型攻击模型,一键生成越狱提示。
  • 对GPT-OSS-20B攻击成功率达81.0%,比直接提示提升54倍。
  • 揭示安全机制在网络安全、欺诈场景中尤为脆弱。

大型语言模型仍易受上下文框架诱导的复杂提示攻击,威胁网络安全应用。本文提出Jailbreak Mimicry,一种系统性方法,通过参数高效微调(LoRA)在Mistral-7B上训练紧凑攻击模型,实现单次输入生成叙事型越狱提示。该方法将手动构造攻击变为可复现的科学流程,用于前瞻性安全评估。基于AdvBench数据集,对GPT-OSS-20B测试集200项实现81.0%攻击成功率(ASR)。跨模型测试显示:对GPT-4为66.5%,对Llama-3为79.5%,对Gemini 2.5 Flash为33.0%,体现广泛适用性与模型差异。相较直接提示(1.5% ASR),提升54倍,暴露当前对齐机制系统性缺陷。分析表明,网络安全(93% ASR)和欺诈类攻击(87.8% ASR)最易突破,物理伤害类则较稳健(55.6% ASR)。采用Claude Sonnet 4自动化评估,并经人工专家验证,保障评估可靠与可扩展性。最后分析失败原因,探讨防御策略。

原文摘要 · Abstract (English)

Large language models (LLMs) remain vulnerable to sophisticated prompt engineering attacks that exploit contextual framing to bypass safety mechanisms, posing significant risks in cybersecurity applications. We introduce Jailbreak Mimicry, a systematic methodology for training compact attacker models to automatically generate narrative-based jailbreak prompts in a one-shot manner. Our approach transforms adversarial prompt discovery from manual craftsmanship into a reproducible scientific process, enabling proactive vulnerability assessment in AI-driven security systems. Developed for the OpenAI GPT-OSS-20B Red-Teaming Challenge, we use parameter-efficient fine-tuning (LoRA) on Mistral-7B with a curated dataset derived from AdvBench, achieving an 81.0% Attack Success Rate (ASR) against GPT-OSS-20B on a held-out test set of 200 items. Cross-model evaluation reveals significant variation in vulnerability patterns: our attacks achieve 66.5% ASR against GPT-4, 79.5% on Llama-3 and 33.0% against Gemini 2.5 Flash, demonstrating both broad applicability and model-specific defensive strengths in cybersecurity contexts. This represents a 54x improvement over direct prompting (1.5% ASR) and demonstrates systematic vulnerabilities in current safety alignment approaches. Our analysis reveals that technical domains (Cybersecurity: 93% ASR) and deception-based attacks (Fraud: 87.8% ASR) are particularly vulnerable, highlighting threats to AI-integrated threat detection, malware analysis, and secure systems, while physical harm categories show greater resistance (55.6% ASR). We employ automated harmfulness evaluation using Claude Sonnet 4, cross-validated with human expert assessment, ensuring reliable and scalable evaluation for cybersecurity red-teaming. Finally, we analyze failure mechanisms and discuss defensive strategies to mitigate these vulnerabilities in AI for cybersecurity.

越狱攻击安全评估提示工程红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。