arXiv:2505.16186cs.AIcs.CL2025-05EMNLP被引 28

通过激活模型内部的‘安全顿悟时刻’,提升对未知攻击的防御能力。

SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning

论文配图:SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
图 1 · 摘自论文原文
  • 在关键句前增强安全信号,提升模型内部表示的安全性。
  • 降低有害响应率9.6%,显著改善对未见越狱攻击的泛化能力。
  • 适合关注大模型安全推理与对抗攻击防护的研究者。

大型推理模型(LRMs)通过显式推理显著提升了复杂任务表现,但面临有害查询和对抗攻击的安全风险。尽管监督微调(SFT)能提升安全性,但其模型在面对未见过的越狱提示时泛化能力差。通过对模型生成过程的深入分析,我们发现安全‘顿悟时刻’通常出现在关键句中,该句反映模型是否将安全地继续回应。基于此,我们提出SafeKey,包含两个互补目标:(1) 双路径安全头,增强关键句前内部表示中的安全信号;(2) 查询掩码建模,强化模型对查询理解的关注,其中包含重要安全线索。在多个安全基准测试中,所提方法显著提升对各类越狱攻击和分布外有害提示的泛化能力,平均有害响应率下降9.6%,同时保持通用能力。分析表明,SafeKey通过重塑内部注意力机制并提升隐藏表示质量来增强安全性能。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) introduce a new generation paradigm of explicitly reasoning before answering, leading to remarkable improvements in complex tasks. However, they pose great safety risks against harmful queries and adversarial attacks. While recent mainstream safety efforts on LRMs, supervised fine-tuning (SFT), improve safety performance, we find that SFT-aligned models struggle to generalize to unseen jailbreak prompts. After thorough investigation of LRMs' generation, we identify a safety aha moment that can activate safety reasoning and lead to a safe response. This aha moment typically appears in the `key sentence', which follows models' query understanding process and can indicate whether the model will proceed safely. Based on these insights, we propose SafeKey, including two complementary objectives to better activate the safety aha moment in the key sentence: (1) a Dual-Path Safety Head to enhance the safety signal in the model's internal representations before the key sentence, and (2) a Query-Mask Modeling objective to improve the models' attention on its query understanding, which has important safety hints. Experiments across multiple safety benchmarks demonstrate that our methods significantly improve safety generalization to a wide range of jailbreak attacks and out-of-distribution harmful prompts, lowering the average harmfulness rate by 9.6\%, while maintaining general abilities. Our analysis reveals how SafeKey enhances safety by reshaping internal attention and improving the quality of hidden representations.

模型安全推理增强越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。