arXiv:2609.08186cs.AIcs.CL2026-09

深度推理会削弱模型安全对齐,提出新指标与防御方法

Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models

论文配图:Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models
图 1 · 摘自论文原文
  • 用注意力稀释解释深度推理导致对齐失效的机制
  • 提出对齐损失率指标,发现推理越深对齐越差
  • 设计轻量级防御法RRA,通过残差连接恢复对齐

链式思维(CoT)为大型推理模型(LRMs)奠定了坚实基础。尽管普遍认为深度推理能提升安全性对齐,但其在长推理过程中的稳定性尚未被充分探索。本文挑战这一观点,揭示了一个关键缺陷:深度推理可能引发对齐崩溃。为此,我们提出对齐损失率(ALR)指标来量化该现象。实验表明,随着推理深度增加,ALR显著上升,说明模型对外部扰动的鲁棒性严重下降。基于此不稳定性,我们提出一种新型越狱范式——推理陷阱(RT),通过诱导模型进入深度推理以放大对抗攻击影响,导致安全能力急剧降低。为进一步揭示崩溃机理,我们识别出注意力稀释是根本原因,源于长推理过程与原始输入之间的注意力竞争。为此,我们提出轻量级防御策略——推理残差对齐(RRA),通过将残差连接融入推理过程,动态强化原始输入信息,从而有效缓解对齐崩溃。

原文摘要 · Abstract (English)

The emergence of Chain-of-Thought (CoT) has established a robust foundation for Large Reasoning Models (LRMs). While deep reasoning is widely believed to enhance safety alignment, the stability of alignment mechanisms under extended reasoning remains underexplored. This paper challenges the prevailing view by revealing a critical vulnerability: Deep Reasoning May Induce Alignment Collapse. To rigorously quantify this phenomenon, we propose the Alignment Loss Rate (ALR) metric. Our experiments demonstrate that as reasoning depth increases, ALR rises significantly, indicating a severe degradation in model robustness against external perturbations. Capitalizing on this instability, a novel jailbreaking paradigm, Reasoning Trap (RT), is proposed. RT induces the model into extended reasoning to amplify the impact of adversarial attacks, leading to a sharp decline in safety capabilities. To elucidate the mechanism behind this collapse, we identify Attention Dilution as the root cause, arising from the competition for attention between the extended reasoning process and the original input. To mitigate this, Reasoning Residual Alignment (RRA), a lightweight defense strategy that dynamically re-emphasizes the input via residual connections integrated with the reasoning process.

推理模型对齐崩溃注意力稀释安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。