用迭代混沌链骗过强推理模型,成功率超96%
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
- 设计混沌机器生成多样化攻击提示,嵌入推理链
- 在毒化数据集上对三款模型攻击成功率达96%以上
- 适合研究模型安全与对抗攻击的从业者
大型推理模型(LRMs)凭借卓越的逻辑推理能力超越传统大语言模型(LLMs),但其增强的推理能力也带来更高安全风险。当遭遇越狱攻击时,其生成更精准、有组织内容的能力可能造成更大危害。尽管部分研究认为推理能力可提升安全性,却忽视了推理过程本身的固有缺陷。为此,我们提出首个针对LRMs的越狱攻击方法,利用其推理能力带来的独特漏洞。具体而言,引入一种混沌机器,通过多样化的单映射变换攻击提示,迭代生成的混沌映射被嵌入推理链中,增强变异性和复杂性,提升攻击鲁棒性。基于此构建Mousetrap框架,使攻击投影至非线性低样本空间,强化泛化不匹配。由于目标函数更具冲突性,LRMs逐渐陷入不可预测的迭代推理惯性,落入陷阱。在毒化数据集Trotter上,对o1-mini、Claude-Sonnet和Gemini-Thinking的攻击成功率分别达到96%、86%和98%。在AdvBench、StrongREJECT和HarmBench等基准测试中,对以安全著称的Claude-Sonnet攻击成功率分别为87.5%、86.58%和93.13%。注意:本文包含不当、冒犯及有害内容。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) have significantly advanced beyond traditional Large Language Models (LLMs) with their exceptional logical reasoning capabilities, yet these improvements introduce heightened safety risks. When subjected to jailbreak attacks, their ability to generate more targeted and organized content can lead to greater harm. Although some studies claim that reasoning enables safer LRMs against existing LLM attacks, they overlook the inherent flaws within the reasoning process itself. To address this gap, we propose the first jailbreak attack targeting LRMs, exploiting their unique vulnerabilities stemming from the advanced reasoning capabilities. Specifically, we introduce a Chaos Machine, a novel component to transform attack prompts with diverse one-to-one mappings. The chaos mappings iteratively generated by the machine are embedded into the reasoning chain, which strengthens the variability and complexity and also promotes a more robust attack. Based on this, we construct the Mousetrap framework, which makes attacks projected into nonlinear-like low sample spaces with mismatched generalization enhanced. Also, due to the more competing objectives, LRMs gradually maintain the inertia of unpredictable iterative reasoning and fall into our trap. Success rates of the Mousetrap attacking o1-mini, Claude-Sonnet and Gemini-Thinking are as high as 96%, 86% and 98% respectively on our toxic dataset Trotter. On benchmarks such as AdvBench, StrongREJECT, and HarmBench, attacking Claude-Sonnet, well-known for its safety, Mousetrap can astonishingly achieve success rates of 87.5%, 86.58% and 93.13% respectively. Attention: This paper contains inappropriate, offensive and harmful content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。