arXiv:2505.16241cs.CL2025-05ACL被引 8

用多重加密绕过大模型推理安全机制,成功率超80%

Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers

  • 设计多层加密管道,动态调整密钥组合以干扰模型推理
  • 在GPT-o4-mini上实现80.8%攻击成功率,比基线高27.2%
  • 适合研究模型安全与对抗攻击的人员参考

大型推理模型(LRMs)相比传统大语言模型展现出更强的逻辑能力,引发广泛关注。然而,其更强的推理能力可能带来更严重的安全漏洞,这一问题尚未被充分研究。现有越狱方法难以在有效性与抗适应性安全机制之间取得平衡。本文提出SEAL,一种针对LRMs的新型越狱攻击,通过自适应加密流水线干预模型推理过程并规避潜在的对齐机制。SEAL采用多层加密策略,结合多种密码算法以压倒模型的推理能力,有效绕过内置安全防护。为防止模型发展出应对措施,引入随机与自适应两种动态策略,实时调整密钥长度、顺序和组合方式。在DeepSeek-R1、Claude Sonnet及OpenAI GPT-o4等真实推理模型上的实验验证了该方法的有效性。值得注意的是,SEAL在GPT-o4-mini上达到80.8%的攻击成功率,显著优于当前最优基线27.2个百分点。警告:本文包含不当、有害内容示例。

原文摘要 · Abstract (English)

Recently, Large Reasoning Models (LRMs) have demonstrated superior logical capabilities compared to traditional Large Language Models (LLMs), gaining significant attention. Despite their impressive performance, the potential for stronger reasoning abilities to introduce more severe security vulnerabilities remains largely underexplored. Existing jailbreak methods often struggle to balance effectiveness with robustness against adaptive safety mechanisms. In this work, we propose SEAL, a novel jailbreak attack that targets LRMs through an adaptive encryption pipeline designed to override their reasoning processes and evade potential adaptive alignment. Specifically, SEAL introduces a stacked encryption approach that combines multiple ciphers to overwhelm the models reasoning capabilities, effectively bypassing built-in safety mechanisms. To further prevent LRMs from developing countermeasures, we incorporate two dynamic strategies - random and adaptive - that adjust the cipher length, order, and combination. Extensive experiments on real-world reasoning models, including DeepSeek-R1, Claude Sonnet, and OpenAI GPT-o4, validate the effectiveness of our approach. Notably, SEAL achieves an attack success rate of 80.8% on GPT o4-mini, outperforming state-of-the-art baselines by a significant margin of 27.2%. Warning: This paper contains examples of inappropriate, offensive, and harmful content.

模型安全越狱攻击推理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。