arXiv:2605.24497cs.AI2026-05

用进化算法动态生成恶意思维链,突破大模型安全防线。

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

论文配图:Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
图 1 · 摘自论文原文
  • 通过角色扮演和碎片化分解,构建可进化的恶意提示池。
  • 自适应变异率与交叉策略提升攻击多样性,超越现有方法。
  • 适合研究模型安全与对抗攻击的开发者参考。

大型推理模型(LRMs)在推理与生成任务中表现出色,但其显式的思维链(CoT)机制引入了新的安全风险,使其易受越狱攻击。现有方法多依赖静态的CoT模板来诱导有害输出,但这类固定设计存在多样性差、适应性弱、效果有限等问题。为此,本文提出一种自适应进化式思维链越狱框架(AE-CoT)。首先,将有害目标改写为温和提示,并通过教师角色扮演分解为语义连贯的推理片段,构建初始的越狱候选池;随后,在结构化表示空间中进行多轮进化搜索,通过片段级交叉与自适应变异率控制机制扩展候选多样性;独立评分模型提供危害性分级评估,高分候选进一步通过有害的CoT模板强化,以诱导更具破坏性的生成结果。在多个模型与数据集上的实验表明,所提方法显著优于现有顶尖越狱方法,具有更强的攻击能力与泛化性。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their explicit chain-of-thought (CoT) mechanism introduces new security risks, making them particularly vulnerable to jailbreak attacks. Existing approaches often rely on static CoT templates to elicit harmful outputs, but such fixed designs suffer from limited diversity, adaptability, and effectiveness. To overcome these limitations, we propose an adaptive evolutionary CoT jailbreak framework, called AE-CoT. Specifically, the method first rewrites harmful goals into mild prompts with teacher role-play and decomposes them into semantically coherent reasoning fragments to construct a pool of CoT jailbreak candidates. Then, within a structured representation space, we perform multi-generation evolutionary search, where candidate diversity is expanded through fragment-level crossover and a mutation strategy with an adaptive mutation-rate control mechanism. An independent scoring model provides graded harmfulness evaluations, and high-scoring candidates are further enhanced with a harmful CoT template to induce more destructive generations. Extensive experiments across multiple models and datasets demonstrate the effectiveness of the proposed AE-CoT, consistently outperforming state-of-the-art jailbreak methods.

模型安全越狱攻击进化算法思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。