通过追溯不安全行为来源,提升推理模型的安全性。
Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model

- 将不安全行为归因于基础模型,动态调整响应优先级。
- 在7个教师模型上平均降低35.4%的攻击成功率。
- 适合关注推理阶段安全性的研究人员和开发者。
尽管拒绝训练在大型语言模型(LLMs)中广泛应用并提升了安全性,但近期研究指出此类对齐方法存在浅层问题。为此,有研究提出通过从更强的推理模型中蒸馏推理能力,实现更深层次的安全对齐。本文研究了这种反思式对齐的效果,发现尽管教师模型更大、安全性更强,但师生模型间仍存在对齐差距,影响学生模型的安全性与通用能力。此外,即使学习了大模型的推理模式,学生模型仍可能保留基础模型的不安全行为。基于此,我们提出一种基于信念的采样(BoN)方法,在隐空间将不安全行为归因于基础模型,从而降低其响应优先级,显著提升安全性,同时保持较低的性能损失。在7个教师模型和6个不同规模的学生模型上,我们在DAN、WildJailbreak和StrongREJECT基准上分别实现28.2%、31.3%和35.4%的攻击成功率下降。进一步实验表明,这些安全提升在强化学习微调后依然有效,揭示了推理安全性的不确定性以及显式归因的重要性。
原文摘要 · Abstract (English)
While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted shortcomings due to the shallow nature of these alignment methods. To this end, the work on Deliberative alignment proposed distilling reasoning capabilities from stronger reasoning models, thereby instilling deeper safety in LLMs. In this work, we study the impact of deliberative alignment in language models. First, we show that despite being larger in model size and stronger in safety capability, there exists an alignment gap between teacher and student language models, which affects both the safety and general utility of the student model. Furthermore, we show that models aligned through deliberative alignment can retain unsafe behaviors from the base model despite learning the reasoning patterns of larger reasoning models. Building upon this observation, we propose a BoN sampling method that attributes the unsafe behavior back to the base LLMs in the latent space, thereby down-ranking unsafe responses to gain a meaningful improvement in model safety across multiple safety benchmarks with minimal loss in utility. In particular, across 7 teacher models and 6 student models of different classes and sizes, we show an average attack success rate (ASR) reduction of 28.2% in DAN, 31.3% in WildJailbreak and 35.4 % in StrongREJECT benchmarks. We further show that these safety gains prevail post RL training, thus highlighting the uncertainty in safety reasoning and it's explicit attribution to the base model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。