用强化学习提升大模型安全回应的精准度,既防风险又不失帮助性。
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

- 基于强化学习构建零样本对齐框架,避免盲目拒绝
- 在多个基准上超越前代模型,安全表现媲美顶级大模型
- 解决误判良性问题的过度防御,适合需要高可信对话场景
大语言模型在各类应用中展现出强大能力,但确保其安全性、有用性和可信度仍是持续挑战。传统以拒绝为主的对齐策略虽能减少有害内容生成,却常因过度拒绝对合法用户需求造成信息屏蔽,无法有效回应敏感请求背后的合理意图。在首创的建设性安全范式Oyster-I基础上,我们识别出其基于监督微调(SFT)方案的两大缺陷:对分布外场景的安全泛化不足,以及安全思维链(CoT)过泛化现象——即安全推理被过度应用于无害查询,降低模型有用性与用户体验。为此,我们提出Oyster-II,一种基于强化学习(RL)的建设性安全对齐框架,采用零强化学习(Zero-RL)范式与多阶段强化学习策略。在广泛基准测试中,Oyster-II全面优于Qwen3-14B及其前代Oyster-I,在安全维度表现接近甚至媲美Qwen3-Max与Qwen3.5-397B。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains a persistent challenge. Conventional refusal-oriented alignment strategies mitigate harmful content generation but systematically fail to serve legitimate user needs, often withholding information that could safely and constructively address the underlying intent of sensitive queries. Building upon the constructive safety paradigm pioneered by Oyster-I, which moves beyond blanket refusal toward thoughtful, response-oriented safety alignment, we identify two critical limitations of its Supervised Fine-Tuning (SFT)-based scheme: insufficient safety generalization to out-of-distribution scenarios and a phenomenon we term safety chain-of-thought (CoT) over-generalization, wherein safety-oriented reasoning patterns are excessively applied to benign queries, degrading helpfulness and user experience. To address these limitations, we propose Oyster-II, a reinforcement learning (RL)-based constructive safety alignment framework that adopts a Zero-RL paradigm combined with a multi-stage reinforcement learning strategy.Evaluated across extensive benchmarks, Oyster-II comprehensively surpasses both Qwen3-14B and its predecessor Oyster-I on safety dimensions, achieving cross-scale performance comparable to Qwen3-Max and Qwen3.5-397B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。