让模型自己生成并内化安全规则,提升对抗攻击下的安全性。
Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models
- 模型从少量正负样本中自动生成安全规则,并自我修正优化。
- 在两个对抗数据集上,安全与不重复拒绝得分提升20.0-45.5分。
- 无需推理时外部规则,内化后的模型仍保持13.3-14.1分增益。
显式安全策略虽可提升推理模型的安全性,但其覆盖范围常落后于不断演化的越狱策略。我们研究了推理模型是否能从少量有害与良性示例中合成并内化特定任务的安全指南。提出自引导自适应安全对齐(SGASA):模型生成指南,基于自身错误进行迭代优化,并通过自评估选择最佳版本,该版本可用于上下文提示或蒸馏至模型中实现无指南推理。在两个对抗提示数据集及三个Qwen3规模模型上,上下文指南使8B与14B模型的综合安全与非重复拒绝得分提升20.0-45.5分。自评估在六组设置中五次选中外部最优优化轮次;指南生成与使用表现出显著不同的缩放模式。仅使用WildJailbreak数据进行对齐监督,内化模型在未启用推理时指南的情况下,仍保留13.3-14.1分的性能增益。结果支持自生成指南作为低资源安全适配的有用中间表示。
原文摘要 · Abstract (English)
Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We study whether a reasoning model can synthesize and internalize a task-specific safety guideline from a small set of harmful and benign examples. We introduce Self-Guided Adaptive Safety Alignment (SGASA). The model generates a guideline, refines it on its own errors, and selects a version by self-evaluation, which can then be applied in context or distilled into the model for guideline-free inference. Across two adversarial prompt datasets and three Qwen3 scales, in-context guidelines improve a combined safety and non-over-refusal score by 20.0-45.5 points on the 8B and 14B models. Self-evaluation selects the externally best refinement round in five of six settings, while guideline generation and utilization show distinct scaling patterns. Using alignment supervision derived only from WildJailbreak, internalized models retain 13.3-14.1 point gains on WildJailbreak without an inference-time guideline. These results support self-generated guidelines as a useful intermediate representation for low-resource safety adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。