arXiv:2504.21038cs.CRcs.AI2025-04被引 9

研究大模型预填充阶段的越狱攻击,发现其成功率超99%且可增强其他攻击。

Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

  • 通过控制模型输出开头实现直接状态篡改,突破传统说服式攻击
  • 自适应方法在多个模型上成功率超99%,显著提升越狱效果
  • 现有内容过滤无效,需关注提示与预填充的异常关联关系

大型语言模型面临越狱攻击的安全威胁。现有研究多聚焦于提示层攻击,而忽视了用户可控响应预填充这一未被充分探索的攻击面。该功能允许攻击者指定模型输出的起始部分,将攻击范式从说服转变为直接状态操纵。本文首次系统性地开展预填充级越狱攻击的黑盒安全分析,对十四种语言模型进行评估。实验表明,预填充攻击成功率极高,自适应方法在多个模型上超过99%。令牌级概率分析显示,攻击通过改变首词概率,使模型从拒绝转为顺从。此外,预填充攻击可作为有效增强器,使现有提示层攻击成功率提升10至15个百分点。对多种防御策略的评估表明,传统内容过滤防护能力有限;基于提示与预填充间操纵关系的检测方法更有效。研究揭示当前大模型安全对齐存在漏洞,亟需在未来的安全训练中关注预填充攻击面。

原文摘要 · Abstract (English)

Large Language Models face security threats from jailbreak attacks. Existing research has predominantly focused on prompt-level attacks while largely ignoring the underexplored attack surface of user-controlled response prefilling. This functionality allows an attacker to dictate the beginning of a model's output, thereby shifting the attack paradigm from persuasion to direct state manipulation.In this paper, we present a systematic black-box security analysis of prefill-level jailbreak attacks. We categorize these new attacks and evaluate their effectiveness across fourteen language models. Our experiments show that prefill-level attacks achieve high success rates, with adaptive methods exceeding 99% on several models. Token-level probability analysis reveals that these attacks work through initial-state manipulation by changing the first-token probability from refusal to compliance.Furthermore, we show that prefill-level jailbreak can act as effective enhancers, increasing the success of existing prompt-level attacks by 10 to 15 percentage points. Our evaluation of several defense strategies indicates that conventional content filters offer limited protection. We find that a detection method focusing on the manipulative relationship between the prompt and the prefill is more effective. Our findings reveal a gap in current LLM safety alignment and highlight the need to address the prefill attack surface in future safety training.

越狱攻击大模型安全预填充黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。