arXiv:2409.00137cs.CRcs.AI2024-09被引 18

提出多轮越狱攻击数据集,揭示攻击结构影响防御效果

Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks

  • 构建可单轮或多轮输入的越狱攻击数据集
  • 多轮攻击成功率显著高于单轮,且防御无效
  • 提醒安全研究需兼顾两种攻击形式

大型语言模型(LLMs)正以惊人速度演进,但其仍易受越狱攻击威胁,且随着模型能力增强,此类攻击日益危险。本文提出一个越狱攻击数据集,其中每个示例均可单轮或多轮输入。我们发现,尽管内容等价,但攻击成功率在两种结构间存在显著差异:针对一种结构的防御无法保证对另一种有效。同样,基于LLM的过滤防护机制也因输入结构不同而表现各异。因此,前沿模型的安全漏洞应同时在单轮和多轮场景下研究;该数据集为此提供了研究工具。

原文摘要 · Abstract (English)

Large language models (LLMs) are improving at an exceptional rate. However, these models are still susceptible to jailbreak attacks, which are becoming increasingly dangerous as models become increasingly powerful. In this work, we introduce a dataset of jailbreaks where each example can be input in both a single or a multi-turn format. We show that while equivalent in content, they are not equivalent in jailbreak success: defending against one structure does not guarantee defense against the other. Similarly, LLM-based filter guardrails also perform differently depending on not just the input content but the input structure. Thus, vulnerabilities of frontier models should be studied in both single and multi-turn settings; this dataset provides a tool to do so.

越狱攻击安全评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。