让大模型先推理安全规则再回答,提升安全性与可靠性。
Deliberative Alignment: Reasoning Enables Safer Language Models
- 训练模型先显式回忆并推理安全规则再生成答案。
- 对齐后模型在抗越狱攻击上更鲁棒,且拒绝率更低。
- 无需人工思维链即可实现可扩展、可解释的对齐。
随着大规模语言模型在安全关键领域的影响日益加深,确保其可靠遵循明确原则仍是核心挑战。我们提出‘审慎对齐’(Deliberative Alignment)新范式,直接教导模型安全规范,并训练其在作答前显式回忆并准确推理这些规范。该方法用于对齐 OpenAI 的 o 系列模型,实现了对 OpenAI 安全政策的高度精确遵循,且无需人类编写的思维链或答案。审慎对齐同时提升了对越狱攻击的鲁棒性并降低了过度拒绝率,还增强了分布外泛化能力。结果表明,对显式规定进行推理,可实现更可扩展、更可信、更可解释的对齐。
原文摘要 · Abstract (English)
As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradigm that directly teaches the model safety specifications and trains it to explicitly recall and accurately reason over the specifications before answering. We used this approach to align OpenAI's o-series models, and achieved highly precise adherence to OpenAI's safety policies, without requiring human-written chain-of-thoughts or answers. Deliberative Alignment pushes the Pareto frontier by simultaneously increasing robustness to jailbreaks while decreasing overrefusal rates, and also improves out-of-distribution generalization. We demonstrate that reasoning over explicitly specified policies enables more scalable, trustworthy, and interpretable alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。