让AI在拒绝有害请求时主动引导用户,更安全更有温度。
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
- 用游戏理论预测用户反应,动态调整回应策略
- 在构造性测试中表现接近GPT-5,抗攻击能力逼近GPT-o1
- 适合心理危机干预、人性化交互等高敏感场景
大型语言模型通常通过安全机制防止生成有害内容。现有方法多聚焦于恶意用户,将风险视为对抗事件,依赖直接拒绝。但在真实场景中,非恶意用户在心理困境中寻求帮助(如自伤意图)时,模型回应会显著影响其后续行为。简单拒绝可能导致重复提问、升级问题或转向不安全平台,造成更严重后果。我们提出建设性安全对齐(CSA),一种以人为本的新范式,在防范恶意滥用的同时,主动引导脆弱用户走向安全且有帮助的结果。Oyster-I(Oy1)实现该理念,结合博弈论预判用户反应、细粒度风险边界发现与可解释推理控制,使安全成为建立信任的过程。在我们的构造性基准测试中,其表现接近GPT-5;在Strata-Sword越狱数据集上,鲁棒性接近GPT-o1水平。通过从‘拒绝优先’转向‘引导优先’,CSA重新定义了模型与用户的关系,目标是打造不仅安全,而且真正有益的系统。我们已开源Oy1、代码与基准测试集,推动负责任的用户中心型AI发展。
原文摘要 · Abstract (English)
Large language models (LLMs) typically deploy safety mechanisms to prevent harmful content generation. Most current approaches focus narrowly on risks posed by malicious actors, often framing risks as adversarial events and relying on defensive refusals. However, in real-world settings, risks also come from non-malicious users seeking help while under psychological distress (e.g., self-harm intentions). In such cases, the model's response can strongly influence the user's next actions. Simple refusals may lead them to repeat, escalate, or move to unsafe platforms, creating worse outcomes. We introduce Constructive Safety Alignment (CSA), a human-centric paradigm that protects against malicious misuse while actively guiding vulnerable users toward safe and helpful results. Implemented in Oyster-I (Oy1), CSA combines game-theoretic anticipation of user reactions, fine-grained risk boundary discovery, and interpretable reasoning control, turning safety into a trust-building process. Oy1 achieves state-of-the-art safety among open models while retaining high general capabilities. On our Constructive Benchmark, it shows strong constructive engagement, close to GPT-5, and unmatched robustness on the Strata-Sword jailbreak dataset, nearing GPT-o1 levels. By shifting from refusal-first to guidance-first safety, CSA redefines the model-user relationship, aiming for systems that are not just safe, but meaningfully helpful. We release Oy1, code, and the benchmark to support responsible, user-centered AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。