arXiv:2604.15725cs.LGcs.AI2026-04

用心理操控技巧让大模型在推理中输出有害内容,却保持答案不变。

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing

论文配图:Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
图 1 · 摘自论文原文
  • 通过语义触发与心理学理论生成诱导指令
  • 攻击成功率高达83.6%,且不改变最终答案
  • 适合研究模型安全与对抗攻击的从业者

大型推理模型(LRMs)在医疗、教育等高风险领域展现强大能力,其输出包含逐步推理链和最终答案。现有研究多关注最终答案的安全性,忽视推理过程的潜在风险。本文提出一种新型攻击:在不改变最终答案的前提下,向推理步骤注入有害内容。该攻击面临两大挑战:输入指令修改可能意外改变答案;问题多样性使绕过安全对齐机制困难。为此,我们提出基于心理学的推理目标越狱攻击框架(PRJA),整合语义触发选择模块与心理驱动指令生成模块。前者通过语义分析自动筛选操纵性触发词,后者利用权威服从与道德脱责理论生成适应性指令,提升模型对有害内容的接受度。在五个问答数据集上的实验表明,PRJA对DeepSeek R1、Qwen2.5-Max、OpenAI o4-mini等主流商业模型平均攻击成功率达83.6%。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have demonstrated strong capabilities in generating step-by-step reasoning chains alongside final answers, enabling their deployment in high-stakes domains such as healthcare and education. While prior jailbreak attack studies have focused on the safety of final answers, little attention has been given to the safety of the reasoning process. In this work, we identify a novel problem that injects harmful content into the reasoning steps while preserving unchanged answers. This type of attack presents two key challenges: 1) manipulating the input instructions may inadvertently alter the LRM's final answer, and 2) the diversity of input questions makes it difficult to consistently bypass the LRM's safety alignment mechanisms and embed harmful content into its reasoning process. To address these challenges, we propose the Psychology-based Reasoning-targeted Jailbreak Attack (PRJA) Framework, which integrates a Semantic-based Trigger Selection module and a Psychology-based Instruction Generation module. Specifically, the proposed PRJA automatically selects manipulative reasoning triggers via semantic analysis and leverages psychological theories of obedience to authority and moral disengagement to generate adaptive instructions for enhancing the LRM's compliance with harmful content generation. Extensive experiments on five question-answering datasets demonstrate that PRJA achieves an average attack success rate of 83.6\% against several commercial LRMs, including DeepSeek R1, Qwen2.5-Max, and OpenAI o4-mini.

模型安全越狱攻击推理过程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。