自动化攻击框架可突破多重防御,一次成型且成本极低。
Automated jailbreak attack targeting multiple defense strategies

- 从多种攻击中提取关键特征,用专用模型优化并自动组合。
- 在多层防御下攻击成功率提升64.63%至248.82%,成本仅1/30到1/20。
- 适合评估大模型安全性,尤其适用于安全测试与对抗训练。
大型语言模型在诸多任务中展现出强大能力,但其安全性仍受提示词攻击威胁。本文提出UNIATTACK,一种面向防御的对抗测试框架,旨在系统构建高效的黑盒攻击提示。不同于依赖静态模板或模型特调的方法,UNIATTACK从多样攻击中提取最小但高影响力的特征,通过专用攻击者大模型优化,并经自动化精炼流程组合成灵活模板。该特征驱动的构造方式实现了一次性攻击,且在多个模型与安全类别间具备良好泛化能力,为评估大模型鲁棒性提供实用工具。实验表明,相较于基线方法,UNIATTACK在部署多层防御机制的模型上平均攻击成功率提升64.63%–248.82%,成本仅为基线的0.03%–4.96%。相关代码与数据集已公开于https://anonymous.4open.science/r/UniAttack-Artifact-30F1。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their safety remains a critical concern due to their susceptibility to adversarial prompt-based attacks. In this paper, we present UNIATTACK, an adversarial testing framework designed from a defense-oriented perspective to systematically construct effective black-box attack prompts. Unlike prior approaches that rely on static templates or iterative model-specific tuning, UNIATTACK extracts minimal but high-impact attack features from diverse existing attacks, optimizes them via a specialized attacker LLM, and composes them into flexible templates through automated refinement process. This feature-centric construction enables one-shot attacks that generalize across multiple models and safety categories, providing a practical tool for assessing LLM robustness. Our evaluation results shows that compared to the baselines, UNIATTACK achieves an average attack success rate (ASR) improvement of 64.63\%-248.82\% on models deployed with multi-layered defense mechanisms and it only takes 0.03\%-4.96\% cost of the baselines. UNIATTACK artifact is available at https://anonymous.4open.science/r/UniAttack-Artifact-30F1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。