通过逐词生成诱骗大模型放弃安全防护,突破拒答机制。
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety

- 分步诱导模型输出单字续写,逐步绕过安全检测。
- 在多个评测集上攻击成功率显著高于现有方法。
- 揭示攻击路径如何抑制拒绝响应的神经表征。
大型语言模型虽经训练可拒绝有害请求,但仍易受基于对话安全机制漏洞的越狱攻击。本文提出增量补全分解(ICD)策略,通过引导模型逐词生成与恶意请求相关的单字续写,最终获取完整响应。我们还设计了利用模型自动生成或攻击者注入中间续写的变体,以及最终响应预填充方案。在多类开源模型上评估表明,ICD在AdvBench、JailbreakBench和StrongREJECT等基准上均优于现有方法。此外,我们提供理论解释并给出机制证据,显示成功攻击路径会抑制与拒绝相关的神经表征,使激活状态偏离安全对齐方向。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to jailbreak attacks that exploit weaknesses in conversational safety mechanisms. We introduce Incremental Completion Decomposition ICD, a trajectory-based jailbreak strategy that elicits a sequence of single-word continuations related to a malicious request before eliciting the full response. In addition, we propose ICD variants that use model-generated or attacker-injected intermediate continuations, as well as final-response prefilling. We evaluate these variants across a broad set of open-weight model families, demonstrating superior Attack Success Rate (ASR) on AdvBench, JailbreakBench, and StrongREJECT compared to existing methods. In addition, we provide a theoretical account of why ICD is effective and present mechanistic evidence that successful attack trajectories suppress refusal-related representations and shift activations away from safety-aligned states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。