arXiv:2601.03662cs.AI2026-01被引 2

通过动态注入安全提醒词,提升大模型推理过程的安全性。

How Does the Thinking Step Influence Model Safety? An Entropy-based Safety Reminder for LRMs

  • 在推理阶段动态插入安全提醒语句,利用熵值触发干预时机。
  • 在六个基准上最高提升45.5%的安全性指标,且不损失推理能力。
  • 适合关注大模型安全、尤其需无参数更新防御方案的研究者。

大推理模型(LRMs)通过显式的思考步骤取得显著成功,但这些思考步骤可能放大不安全行为,带来新风险。现有防御方法因忽略LRM独特的推理动态而失效。本文发现,思考步骤中出现安全提醒语句对保障模型安全至关重要。为此,我们提出SafeRemind——一种解码时防御方法,动态向思考步骤注入安全提醒语句。通过熵值触发机制,在决策锁定点进行干预,引导潜在有害路径走向更安全结果,无需任何参数更新。在五个LRMs和六个基准上的广泛评估表明,SafeRemind显著提升安全性,最高提升达45.5%p,同时保持核心推理能力。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) achieve remarkable success through explicit thinking steps, yet the thinking steps introduce a novel risk by potentially amplifying unsafe behaviors. Despite this vulnerability, conventional defense mechanisms remain ineffective as they overlook the unique reasoning dynamics of LRMs. In this work, we find that the emergence of safe-reminding phrases within thinking steps plays a pivotal role in ensuring LRM safety. Motivated by this finding, we propose SafeRemind, a decoding-time defense method that dynamically injects safe-reminding phrases into thinking steps. By leveraging entropy triggers to intervene at decision-locking points, SafeRemind redirects potentially harmful trajectories toward safer outcomes without requiring any parameter updates. Extensive evaluations across five LRMs and six benchmarks demonstrate that SafeRemind substantially enhances safety, achieving improvements of up to 45.5%p while preserving core reasoning utility.

大模型安全推理增强防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。