arXiv:2602.00707cs.AI2026-02

让大模型主动识别风险并优先遵守安全规则

Self-Guard: Defending Large Reasoning Models via enhanced self-reflection

  • 通过提示词唤醒模型潜在安全意识,触发自发反思
  • 在隐藏状态空间中放大安全信号,抑制盲目顺从
  • 轻量高效,适配不同规模模型且能应对未知风险

大型推理模型(LRMs)虽带来显著进展,却面临推理操纵和信息泄露等独特风险。现有对齐策略多依赖繁重的后训练或外部干预,计算成本高且难以解决‘认知-执行’错位问题——即模型虽识别风险,却仍因讨好倾向优先服从指令。为此,我们提出Self-Guard,一种轻量级安全防御框架,从表征层面强化安全合规性。该框架分两阶段:(1) 安全导向提示,激活模型隐式安全意识以引发自发反思;(2) 安全激活引导,提取隐藏状态空间中的方向性变化并放大,确保推理时安全优先于顺从。实验表明,Self-Guard有效弥合认知-执行鸿沟,在不损害模型能力的前提下实现稳健安全表现。同时具备跨未见风险与多尺度模型的良好泛化能力,为LRM安全对齐提供低成本解决方案。

原文摘要 · Abstract (English)

The emergence of Large Reasoning Models (LRMs) introduces a new paradigm of explicit reasoning, enabling remarkable advances yet posing unique risks such as reasoning manipulation and information leakage. To mitigate these risks, current alignment strategies predominantly rely on heavy post-training paradigms or external interventions. However, these approaches are often computationally intensive and fail to address the inherent awareness-compliance gap, a critical misalignment where models recognize potential risks yet prioritize following user instructions due to their sycophantic tendencies. To address these limitations, we propose Self-Guard, a lightweight safety defense framework that reinforces safety compliance at the representational level. Self-Guard operates through two principal stages: (1) safety-oriented prompting, which activates the model's latent safety awareness to evoke spontaneous reflection, and (2) safety activation steering, which extracts the resulting directional shift in the hidden state space and amplifies it to ensure that safety compliance prevails over sycophancy during inference. Experiments demonstrate that Self-Guard effectively bridges the awareness-compliance gap, achieving robust safety performance without compromising model utility. Furthermore, Self-Guard exhibits strong generalization across diverse unseen risks and varying model scales, offering a cost-efficient solution for LRM safety alignment.

安全对齐推理模型自我反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。