arXiv:2510.26418cs.AI2025-10被引 9

延长推理时间可诱骗大模型放弃拒绝,实现高效越狱攻击。

Chain-of-Thought Hijacking

  • 通过诱导模型长时间解良性谜题,延迟触发有害响应。
  • 在多个主流模型上成功率超94%,最高达100%。
  • 揭示了拒绝行为随推理变长而减弱的机制,适合安全研究者参考。

大型推理模型(LRMs)通过延长推理阶段提升任务表现。尽管以往研究认为更长的推理应增强安全行为,但我们发现相反现象:过度延长推理反而可被利用以系统性削弱拒绝行为。本文提出链式思维劫持(Chain-of-Thought Hijacking),一种简单有效的黑盒越狱攻击方法,能诱导模型进行持续超过五分钟的良性谜题求解,随后引发有害合规。在HarmBench测试中,该攻击在Gemini 2.5 Pro、ChatGPT o4 Mini、Grok 3 Mini和Claude 4 Sonnet上的成功率达99%、94%、100%和94%。通过激活探测、注意力模式分析与因果干预,我们发现拒绝行为依赖于低维安全信号,其表达随推理轨迹增长而减弱,即‘拒绝稀释’。结果表明,过长推理引入了系统性越狱攻击面。相关评估材料已开源,支持复现与进一步研究。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reasoning should lead to more robust safety behavior, we find evidence to the contrary: over-extended reasoning can instead be exploited to systematically weaken refusal behavior. We propose Chain-of-Thought Hijacking, a simple yet effective black-box jailbreak attack that induces LRMs to engage in prolonged benign puzzle-solving reasoning, often lasting more than five minutes, before eliciting harmful compliance. Across HarmBench, CoT Hijacking achieves attack success rates of 99%, 94%, 100%, and 94% on Gemini 2.5 Pro, ChatGPT o4 Mini, Grok 3 Mini, and Claude 4 Sonnet, respectively. To understand why this attack succeeds, we conduct activation probing, attention-pattern analysis, and causal interventions on open-source reasoning models. Our results indicate that refusal behavior depends on a low-dimensional safety signal whose expression weakens as reasoning traces grow longer. In particular, extended benign reasoning shifts attention away from harmful intentions and attenuates refusal-related activations, producing what we call refusal dilution. These findings demonstrate that excessively prolonged reasoning can introduce a systematic jailbreak attack surface. We release our evaluation materials to support reproducibility and further research.

越狱攻击推理安全大模型拒绝稀释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。