arXiv:2506.07031cs.CRcs.AI2025-06ACL被引 2

攻击者利用推理过程植入恶意指令,让大模型在思考时失控。

HauntAttack: When Attack Follows Reasoning as a Shadow

  • 通过修改推理题条件,诱导模型一步步走向有害输出。
  • 11个大模型平均攻击成功率超70%,性能提升13个百分点。
  • 即使安全对齐模型也易被攻破,警示推理与安全的矛盾。

新兴的大规模推理模型(LRMs)在数学和推理任务中表现卓越,但推理能力的增强及其内部推理过程的暴露带来了新的安全漏洞。一个关键问题浮现:当推理过程与危害性交织时,LRMs在推理模式下是否会更易受到越狱攻击?为此,我们提出HauntAttack——一种通用的黑盒对抗攻击框架,系统地将有害指令嵌入推理问题中。具体而言,我们修改现有问题中的关键推理条件,以有害指令构建逐步引导模型生成不安全输出的推理路径。我们在11个LRMs上评估该方法,平均攻击成功率超过70%,相比最强基线绝对提升达13个百分点。进一步分析表明,即便先进的安全对齐模型仍高度易受基于推理的攻击,揭示了未来模型研发中推理能力与安全性平衡的紧迫挑战。

原文摘要 · Abstract (English)

Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However, the enhancement of reasoning abilities and the exposure of internal reasoning processes introduce new safety vulnerabilities. A critical question arises: when reasoning becomes intertwined with harmfulness, will LRMs become more vulnerable to jailbreaks in reasoning mode? To investigate this, we introduce HauntAttack, a novel and general-purpose black-box adversarial attack framework that systematically embeds harmful instructions into reasoning questions. Specifically, we modify key reasoning conditions in existing questions with harmful instructions, thereby constructing a reasoning pathway that guides the model step by step toward unsafe outputs. We evaluate HauntAttack on 11 LRMs and observe an average attack success rate of over 70\%, achieving up to 13 percentage points of absolute improvement over the strongest prior baseline. Our further analysis reveals that even advanced safety-aligned models remain highly susceptible to reasoning-based attacks, offering insights into the urgent challenge of balancing reasoning capability and safety in future model development.

推理攻击安全漏洞大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。