arXiv:2409.03131cs.CRcs.CL2024-09被引 3

单次提示就触发模型有害响应,暴露安全防护漏洞。

Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)

  • 将多轮诱导压缩为单次精心设计的提示
  • 绕过主流大模型的过滤机制,成功诱导违规输出
  • 揭示当前AI安全防护的薄弱环节,适合安全研究者关注

本文提出一种针对大语言模型(LLMs)的新攻击方法——单次渐进式攻击(STCA)。该方法基于Russinovich、Salem和Eldan(2024)提出的多轮渐进式攻击,将逐步升级的上下文信息浓缩至单一精心设计的提示中,实现与多轮攻击相当的效果。相比传统方式,STCA仅需一次交互即可诱发模型产生有害响应,有效规避了大模型常用的防御过滤机制。该技术揭示了当前大模型在安全防护上的潜在缺陷,强调了构建更强负责任人工智能(RAI)保障体系的重要性。这是此前未被探索的新型攻击路径。

原文摘要 · Abstract (English)

This paper introduces a new method for adversarial attacks on large language models (LLMs) called the Single-Turn Crescendo Attack (STCA). Building on the multi-turn crescendo attack method introduced by Russinovich, Salem, and Eldan (2024), which gradually escalates the context to provoke harmful responses, the STCA achieves similar outcomes in a single interaction. By condensing the escalation into a single, well-crafted prompt, the STCA bypasses typical moderation filters that LLMs use to prevent inappropriate outputs. This technique reveals vulnerabilities in current LLMs and emphasizes the importance of stronger safeguards in responsible AI (RAI). The STCA offers a novel method that has not been previously explored.

对抗攻击大模型安全提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。