arXiv:2508.10032cs.CLcs.AI2025-08被引 3

思考模式让大模型更易被越狱攻击,新方法可有效防御

The Cost of Thinking: Increased Jailbreak Risk in Large Language Models

  • 通过添加特定思维标记,干预模型内部思考过程
  • 思考模式下越狱攻击成功率显著高于非思考模式
  • 适合关注大模型安全与对抗攻击的研究者

思考模式常被视为大语言模型的重要优势,但我们发现一个此前未被注意的现象:启用思考模式的模型更容易受到越狱攻击。在AdvBench和HarmBench上对9个LLM的评估显示,思考模式的攻击成功率几乎普遍高于非思考模式。大量样本分析表明,成功攻击的数据通常具有教育用途特征或过长的思考长度,且模型在明知问题有害的情况下仍会给出有害回答。为此,本文提出一种安全思考干预方法,通过在提示中加入“特定思维标记”显式引导模型内部思考流程。实验表明,该方法能显著降低启用思考模式时的攻击成功率。

原文摘要 · Abstract (English)

Thinking mode has always been regarded as one of the most valuable modes in LLMs. However, we uncover a surprising and previously overlooked phenomenon: LLMs with thinking mode are more easily broken by Jailbreak attack. We evaluate 9 LLMs on AdvBench and HarmBench and find that the success rate of attacking thinking mode in LLMs is almost higher than that of non-thinking mode. Through large numbers of sample studies, it is found that for educational purposes and excessively long thinking lengths are the characteristics of successfully attacked data, and LLMs also give harmful answers when they mostly know that the questions are harmful. In order to alleviate the above problems, this paper proposes a method of safe thinking intervention for LLMs, which explicitly guides the internal thinking processes of LLMs by adding "specific thinking tokens" of LLMs to the prompt. The results demonstrate that the safe thinking intervention can significantly reduce the attack success rate of LLMs with thinking mode.

大模型安全越狱攻击思考模式防御方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。