arXiv:2502.05223cs.CRcs.AI2025-02被引 5

KDA模型可自动生成多样攻击提示,高效突破大模型安全防护。

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

  • 将多个顶尖攻击者知识融合为单一模型,自动生成攻击提示。
  • 在多款开源与商业模型上成功率更高,耗时减少50%以上。
  • 适合安全测试人员快速评估大模型防御能力。

越狱攻击通过特定提示绕过大语言模型的安全机制,导致模型生成有害、不适当及偏差内容。现有方法严重依赖精心设计的系统提示和大量查询才能实现单次攻击,成本高且难以大规模应用。为此,我们提出将一组先进攻击者的知识蒸馏为一个开源模型——知识蒸馏攻击者(KDA),该模型经微调后可无需复杂系统提示工程,自动生成连贯且多样化的攻击提示。相比现有攻击方法,KDA在针对多个先进开源与商业黑箱大模型时展现出更高的攻击成功率和显著的成本-时间效率。此外,我们对基线方法与KDA生成的提示进行了定量多样性分析,发现提示多样性及集成攻击是其高效性的关键因素。

原文摘要 · Abstract (English)

Jailbreak attacks exploit specific prompts to bypass LLM safeguards, causing the LLM to generate harmful, inappropriate, and misaligned content. Current jailbreaking methods rely heavily on carefully designed system prompts and numerous queries to achieve a single successful attack, which is costly and impractical for large-scale red-teaming. To address this challenge, we propose to distill the knowledge of an ensemble of SOTA attackers into a single open-source model, called Knowledge-Distilled Attacker (KDA), which is finetuned to automatically generate coherent and diverse attack prompts without the need for meticulous system prompt engineering. Compared to existing attackers, KDA achieves higher attack success rates and greater cost-time efficiency when targeting multiple SOTA open-source and commercial black-box LLMs. Furthermore, we conducted a quantitative diversity analysis of prompts generated by baseline methods and KDA, identifying diverse and ensemble attacks as key factors behind KDA's effectiveness and efficiency.

越狱攻击提示生成安全评测模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。