通过语义湍流检测大模型越狱攻击,无需额外模型。
The Laminar Flow Hypothesis: Detecting Jailbreaks via Semantic Turbulence in Large Language Models
- 用层间余弦速度方差衡量模型推理时的平滑性,识别异常轨迹。
- 攻击下Qwen2-1.5B的语义湍流提升75.4%,显著高于正常输入。
- 可实时检测越狱并揭示模型安全机制类型,适合安全研究者。
随着大语言模型(LLMs)广泛应用,抵御对抗性越狱攻击的挑战日益严峻。现有防御策略多依赖计算开销大的外部分类器或脆弱的词汇过滤器,忽视了模型内在推理过程的动力学特性。本文提出层流假设:良性输入在高维隐空间中引发平滑、渐进的转换,而越狱提示则触发混沌、高方差的轨迹——称为语义湍流,源于安全对齐与指令遵循目标间的内部冲突。为此提出一种零样本度量:层间余弦速度方差。在多种小型语言模型上的实验表明,该指标具有显著诊断能力。经RLHF对齐的Qwen2-1.5B在攻击下湍流值上升75.4%(p < 0.001),验证了内部冲突假说;而Gemma-2B则表现出22.0%的湍流下降,呈现独特的低熵‘反射式’拒绝机制。结果表明,语义湍流不仅可作为轻量级实时越狱检测器,还可非侵入式诊断黑盒模型的安全架构。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become ubiquitous, the challenge of securing them against adversarial "jailbreaking" attacks has intensified. Current defense strategies often rely on computationally expensive external classifiers or brittle lexical filters, overlooking the intrinsic dynamics of the model's reasoning process. In this work, the Laminar Flow Hypothesis is introduced, which posits that benign inputs induce smooth, gradual transitions in an LLM's high-dimensional latent space, whereas adversarial prompts trigger chaotic, high-variance trajectories - termed Semantic Turbulence - resulting from the internal conflict between safety alignment and instruction-following objectives. This phenomenon is formalized through a novel, zero-shot metric: the variance of layer-wise cosine velocity. Experimental evaluation across diverse small language models reveals a striking diagnostic capability. The RLHF-aligned Qwen2-1.5B exhibits a statistically significant 75.4% increase in turbulence under attack (p less than 0.001), validating the hypothesis of internal conflict. Conversely, Gemma-2B displays a 22.0% decrease in turbulence, characterizing a distinct, low-entropy "reflex-based" refusal mechanism. These findings demonstrate that Semantic Turbulence serves not only as a lightweight, real-time jailbreak detector but also as a non-invasive diagnostic tool for categorizing the underlying safety architecture of black-box models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。