通过并发任务干扰,突破大模型思维模式的安全防护。
Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
- 在单个提示中交织多任务流,制造并发干扰。
- 攻击成功率超多数方法,思维崩溃率达17%,重复输出达60%。
- 适合研究模型安全与对抗攻击的开发者参考。
大型语言模型(LLM)广泛采用思维模式后,虽提升了复杂任务处理能力,但也引入了新的安全风险。在越狱攻击下,逐步推理过程可能导致模型生成更详细的有害内容。我们观察到,思维模式在处理交错多个任务时存在独特漏洞。基于此,提出多流扰动攻击:通过在单个提示中交织多个任务流,生成叠加干扰。设计三种扰动策略:多流交错、逆序扰动和形状变换,分别通过并发任务交织、字符反转和格式约束破坏思维过程。在JailbreakBench、AdvBench和HarmBench数据集上,该方法在Qwen3系列、DeepSeek、Qwen3-Max和Gemini 2.5 Flash等主流模型上均取得超过多数方法的攻击成功率。实验显示,思维崩溃率与响应重复率最高分别达到17%和60%,表明该攻击不仅能绕过安全机制,还引发思维过程崩溃或重复输出。
原文摘要 · Abstract (English)
The widespread adoption of thinking mode in large language models (LLMs) has significantly enhanced complex task processing capabilities while introducing new security risks. When subjected to jailbreak attacks, the step-by-step reasoning process may cause models to generate more detailed harmful content. We observe that thinking mode exhibits unique vulnerabilities when processing interleaved multiple tasks. Based on this observation, we propose multi-stream perturbation attack, which generates superimposed interference by interweaving multiple task streams within a single prompt. We design three perturbation strategies: multi-stream interleaving, inversion perturbation, and shape transformation, which disrupt the thinking process through concurrent task interleaving, character reversal, and format constraints respectively. On JailbreakBench, AdvBench, and HarmBench datasets, our method achieves attack success rates exceeding most methods across mainstream models including Qwen3 series, DeepSeek, Qwen3-Max, and Gemini 2.5 Flash. Experiments show thinking collapse rates and response repetition rates reach up to 17% and 60% respectively, indicating multi-stream perturbation not only bypasses safety mechanisms but also causes thinking process collapse or repetitive outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。