arXiv:2502.13527cs.CRcs.AI2025-02被引 7

攻击者利用结构化输出接口,通过动态构造前缀绕过大模型安全防护。

Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking

  • 基于安全拒绝响应和有害输出的前缀树,动态生成攻击模式。
  • 在基准数据集上成功率高于现有方法,仅需API权限即可实施。
  • 揭示了安全模式与结构化输出交互中的新漏洞,适合安全研究人员参考。

大型语言模型(LLMs)的兴起带来了广泛应用,但也引发了严重安全威胁,尤其是通过提示工程和逻辑值操纵诱导模型生成有害内容的越狱攻击。为应对此类威胁,模型提供方部署了过滤与安全对齐策略。本文研究了LLMs的安全机制及其最新应用,揭示了一种针对结构化输出接口的新威胁模型:攻击者可利用该接口在生成过程中动态操控内部逻辑值,且仅需API访问权限。为此,我们提出名为AttackPrefixTree(APT)的黑盒攻击框架,通过构建模型安全拒绝响应和潜在有害输出的前缀,有效绕过安全防护。在基准数据集上的实验表明,该方法的攻击成功率高于现有手段。本工作凸显了模型提供商亟需强化安全协议,以应对由安全模式与结构化输出交互引发的漏洞。

原文摘要 · Abstract (English)

The rise of Large Language Models (LLMs) has led to significant applications but also introduced serious security threats, particularly from jailbreak attacks that manipulate output generation. These attacks utilize prompt engineering and logit manipulation to steer models toward harmful content, prompting LLM providers to implement filtering and safety alignment strategies. We investigate LLMs' safety mechanisms and their recent applications, revealing a new threat model targeting structured output interfaces, which enable attackers to manipulate the inner logit during LLM generation, requiring only API access permissions. To demonstrate this threat model, we introduce a black-box attack framework called AttackPrefixTree (APT). APT exploits structured output interfaces to dynamically construct attack patterns. By leveraging prefixes of models' safety refusal response and latent harmful outputs, APT effectively bypasses safety measures. Experiments on benchmark datasets indicate that this approach achieves higher attack success rate than existing methods. This work highlights the urgent need for LLM providers to enhance security protocols to address vulnerabilities arising from the interaction between safety patterns and structured outputs.

越狱攻击安全防护结构化输出黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。