让开源大模型更安全,且不损失推理能力
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
- 用1.5万条安全指令生成的推理轨迹训练模型
- 在对抗恶意请求和越狱攻击时表现显著提升
- 保持原始推理能力,适合需安全性的实际应用
大型推理模型(LRMs)如OpenAI o1和DeepSeek-R1在数学与编程等复杂任务中表现突破。然而,开源版DeepSeek-R1存在安全风险,易响应恶意请求,影响实际应用。本文提出RealSafe-R1,为DeepSeek-R1设计的安全对齐版本。通过构建1.5万条在明确拒绝指令下生成的安全感知推理轨迹数据集进行训练。定量实验与定性案例均显示,模型在抵御有害请求和越狱攻击方面显著增强。重要的是,本方法通过维持生成数据分布不变,避免了以往安全对齐导致推理能力下降的问题,完整保留原模型推理性能。RealSafe-R1模型权重已开源,可在HuggingFace获取。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs), such as OpenAI o1 and DeepSeek-R1, have been rapidly progressing and achieving breakthrough performance on complex reasoning tasks such as mathematics and coding. However, the open-source R1 models have raised safety concerns in wide applications, such as the tendency to comply with malicious queries, which greatly impacts the utility of these powerful models in their applications. In this paper, we introduce RealSafe-R1 as safety-aligned versions of DeepSeek-R1 distilled models. To train these models, we construct a dataset of 15k safety-aware reasoning trajectories generated by DeepSeek-R1, under explicit instructions for expected refusal behavior. Both quantitative experiments and qualitative case studies demonstrate the models' improvements, which are shown in their safety guardrails against both harmful queries and jailbreak attacks. Importantly, unlike prior safety alignment efforts that often compromise reasoning performance, our method preserves the models' reasoning capabilities by maintaining the training data within the original distribution of generation. Model weights of RealSafe-R1 are open-source at https://huggingface.co/RealSafe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。