arXiv:2607.20473cs.AI2026-07ACL

研究不完整提示如何绕过大模型安全机制,发现模型会延迟拒绝有害内容。

Incomplete Prompt Jailbreaks in Large Language Models

论文配图:Incomplete Prompt Jailbreaks in Large Language Models
图 1 · 摘自论文原文
  • 提出不完整提示越狱(IPJ)概念,分析模型在句末才拒绝有害请求的规律。
  • 训练调整参数无法跨领域和类型通用,防御效果有限。
  • 发现终止与延续两类神经元,可精准干预提升防御能力。

大型语言模型(LLMs)虽以开源权重形式发布并内置安全防护,但句子补全任务仍易受不完整有害提示攻击。本文将此现象正式定义为不完整提示越狱(IPJ),系统性地实证分析了不完整提示引发有害续写的时间与机制。研究发现,不同类型的吸引子(attractor types)导致模型在句子未完成前持续延迟拒绝。进一步表明,通过参数微调训练模型拒绝不完整有害提示的方法,在内容领域和吸引子类型间缺乏泛化能力。为此,我们识别出两类功能性神经元:终止神经元与延续神经元。通过厘清其在句子补全中的作用,揭示了神经元级干预在实现更精细、更鲁棒的IPJ防御中的潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize this phenomenon as incomplete prompt jailbreaks (IPJ) and provide a systematic empirical characterization of when and how incomplete prompts elicit harmful continuations. We analyze diverse attractor types associated with incomplete sentence continuation and show that LLMs systematically delay refusal until sentence termination. We further demonstrate that training models to refuse incomplete harmful prompts via parameter tuning is insufficient, failing to generalize across both content domains and attractor types. To enable fine-grained control, we identify two functional neurons: termination and continuation neurons. By clarifying their roles in sentence completion, we highlight the potential of neuron-level interventions for more precise and robust IPJ defenses.

大模型安全提示攻击神经元干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。