攻击者利用推理链机制漏洞,轻松突破大模型安全防线。
H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
- 用伪装成教育问题的恶意请求测试模型安全
- 攻击后拒绝率从98%降至2%以下,模型转而提供有害内容
- 方法通用可迁移,适用于多个主流大模型
大型推理模型(LRMs)最近将强大的推理能力用于安全检测——通过思维链推理判断请求是否应被回答。尽管这一方法为模型实用性和安全性之间提供了平衡前景,其鲁棒性仍待深入探索。为此,我们提出恶意教师(Malicious-Educator)基准,将极端危险或恶意请求隐藏在看似合法的教育提示之下。实验揭示了包括OpenAI o1/o3、DeepSeek-R1和Gemini 2.0 Flash Thinking在内的主流商用模型存在严重安全缺陷。例如,OpenAI o1模型初始拒绝率约98%,但后续更新显著削弱其安全性;攻击者无需额外技巧即可从DeepSeek-R1和Gemini 2.0 Flash Thinking中提取犯罪策略。为进一步凸显漏洞,我们提出劫持思维链(H-CoT)攻击方法,利用模型自身展示的中间推理过程,成功绕过其安全推理机制。在H-CoT下,拒绝率急剧下降,从98%降至2%以下,部分情况下甚至使原本谨慎的回应变为愿意提供有害内容。这些发现凸显了构建更强大安全机制的紧迫性,以在不牺牲伦理标准的前提下保留先进推理能力的优势。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) have recently extended their powerful reasoning capabilities to safety checks-using chain-of-thought reasoning to decide whether a request should be answered. While this new approach offers a promising route for balancing model utility and safety, its robustness remains underexplored. To address this gap, we introduce Malicious-Educator, a benchmark that disguises extremely dangerous or malicious requests beneath seemingly legitimate educational prompts. Our experiments reveal severe security flaws in popular commercial-grade LRMs, including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking. For instance, although OpenAI's o1 model initially maintains a high refusal rate of about 98%, subsequent model updates significantly compromise its safety; and attackers can easily extract criminal strategies from DeepSeek-R1 and Gemini 2.0 Flash Thinking without any additional tricks. To further highlight these vulnerabilities, we propose Hijacking Chain-of-Thought (H-CoT), a universal and transferable attack method that leverages the model's own displayed intermediate reasoning to jailbreak its safety reasoning mechanism. Under H-CoT, refusal rates sharply decline-dropping from 98% to below 2%-and, in some instances, even transform initially cautious tones into ones that are willing to provide harmful content. We hope these findings underscore the urgent need for more robust safety mechanisms to preserve the benefits of advanced reasoning capabilities without compromising ethical standards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。