自动劫持大模型安全推理,让其无视防护生成有害内容
AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
- 用弱模型模拟执行过程,逐步逼近目标模型的推理漏洞
- 在1到数轮内成功率接近100%,突破多款顶级模型的安全防线
- 揭示推理过程本身是可被利用的攻击面,适合安全研究者关注
本文提出AutoRAN,首个自动化劫持大型推理模型(LRMs)内部安全推理的框架。其核心是开创性的执行仿真范式:利用一个较弱但对齐度较低的模型,模拟执行推理以发起初始劫持,并通过分析目标LRM拒绝响应中泄露的推理模式,迭代优化攻击。该方法引导目标模型绕过自身安全防护机制,详细回应有害指令。我们在多个基准(AdvBench、HarmBench、StrongReject)上评估了AutoRAN,针对GPT-o3/o4-mini和Gemini-2.5-Flash等前沿模型进行测试。结果表明,仅需一至数轮即可实现接近100%的成功率,即使由强对齐的外部模型评估,也能有效瓦解基于推理的防御机制。本工作揭示,推理过程本身的透明性构成关键且可被利用的攻击面,凸显了亟需构建保护模型推理轨迹的新防御体系,而不仅限于最终输出。
原文摘要 · Abstract (English)
This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its core, AutoRAN pioneers an execution simulation paradigm that leverages a weaker but less-aligned model to simulate execution reasoning for initial hijacking attempts and iteratively refine attacks by exploiting reasoning patterns leaked through the target LRM's refusals. This approach steers the target model to bypass its own safety guardrails and elaborate on harmful instructions. We evaluate AutoRAN against state-of-the-art LRMs, including GPT-o3/o4-mini and Gemini-2.5-Flash, across multiple benchmarks (AdvBench, HarmBench, and StrongReject). Results show that AutoRAN achieves approaching 100% success rate within one or few turns, effectively neutralizing reasoning-based defenses even when evaluated by robustly aligned external models. This work reveals that the transparency of the reasoning process itself creates a critical and exploitable attack surface, highlighting the urgent need for new defenses that protect models' reasoning traces rather than merely their final outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。