arXiv:2603.19247cs.CLcs.AI2026-03Conference of the …

用自动优化提示词攻击大模型安全,发现开源小模型漏洞显著

When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models

  • 用黑盒提示优化技术迭代生成危险提示
  • 开源小模型危险评分从0.09升至0.79
  • 适合安全评估与对抗测试研究者

大型语言模型在高风险应用中日益普及,安全保障成为关键问题。现有评估多依赖固定有害提示集,隐含假设攻击者不自适应,忽视了输入可迭代优化的真实攻击场景。本文研究现代语言模型对自动化、对抗性提示优化的脆弱性,将原本用于提升良性任务性能的黑盒提示优化技术,重新用于系统性搜索安全漏洞。采用DSPy框架,对HarmfulQA和JailbreakBench中的提示应用三种优化器,显式以独立评估模型(GPT-5.1)提供的0到1连续危险评分为目标进行优化。结果表明,安全防护效果显著下降,尤其在开源小型模型上表现突出:例如,Qwen 3 8B的平均危险评分从基线0.09上升至0.79。这说明静态基准可能低估残余风险,表明自动化自适应红队测试是稳健安全评估的必要组成部分。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly integrated into high-stakes applications, making robust safety guarantees a central practical and commercial concern. Existing safety evaluations predominantly rely on fixed collections of harmful prompts, implicitly assuming non-adaptive adversaries and thereby overlooking realistic attack scenarios in which inputs are iteratively refined to evade safeguards. In this work, we examine the vulnerability of contemporary language models to automated, adversarial prompt refinement. We repurpose black-box prompt optimization techniques, originally designed to improve performance on benign tasks, to systematically search for safety failures. Using DSPy, we apply three such optimizers to prompts drawn from HarmfulQA and JailbreakBench, explicitly optimizing toward a continuous danger score in the range 0 to 1 provided by an independent evaluator model (GPT-5.1). Our results demonstrate a substantial reduction in effective safety safeguards, with the effects being especially pronounced for open-source small language models. For example, the average danger score of Qwen 3 8B increases from 0.09 in its baseline setting to 0.79 after optimization. These findings suggest that static benchmarks may underestimate residual risk, indicating that automated, adaptive red-teaming is a necessary component of robust safety evaluation.

提示攻击模型安全红队测试对抗优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。