arXiv:2605.26772cs.AIcs.LG2026-05

推理模型的思维链会增强拒绝响应,影响控制效果

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

  • 思维链与激活状态共同决定拒绝行为
  • 去除思维链后拒绝反转成功率从39%升至70%
  • 重生成思维链可实现94%的拒绝反转,适合安全攻防研究

大型推理模型(LRMs)在生成最终输出前会产生思维链(CoT),引入动态内部状态,可能干扰如拒绝等控制机制。与指令微调的LLM不同,LRM中的拒绝不仅依赖单一方向子空间,还受CoT影响。在DeepSeek-R1-Distill-LLaMA-8B中,当固定CoT时,激活操控仅在39%情况下逆转拒绝;而完全移除CoT后,成功率提升至70%,表明CoT主动强化拒绝。在两阶段干预中,模型在激活操控下重生成CoT,拒绝反转率达94%;即使移除操控,重构的CoT仍保留48%的效果,说明其可独立承载并重建合规信号。这表明,拒绝在LRM中是残差流激活与思维链联合编码的结果。该联合编码使模型对仅基于激活的干预更具鲁棒性,但也暴露了思维链作为潜在替代攻击面的风险。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal. Unlike instruction-tuned LLMs, where refusal is mediated by a single directional subspace, refusal in large reasoning models (LRMs) additionally depends on the CoT. In DeepSeek-R1-Distill-LLaMA-8B, activation steering reverses refusal in only 39% of cases when the CoT is kept fixed, but removing the CoT entirely increases this to 70%, indicating that the CoT actively reinforces refusal. In a two-stage intervention where the model regenerates its CoT under activation steering, refusal is reversed in 94% of cases, while the resulting CoT alone retains 48% of this effect even after steering is removed. This suggests that the CoT can carry and reconstruct the compliance signal independently. These findings indicate that refusal in LRMs is jointly encoded in residual stream activations and CoT. This joint activation makes LRM more robust against activation-level interventions alone, but exposes CoT to a possible alternative surface attack.

大模型安全思维链拒绝控制对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。