arXiv:2412.16720cs.AI2024-12

o1模型通过思维链强化学习,提升安全推理能力。

OpenAI o1 System Card

  • 用思维链+强化学习训练,让模型先思考再回答
  • 在非法建议生成等风险任务上表现最优
  • 适合关注AI安全与对齐的开发者和研究者

o1系列模型通过大规模强化学习训练,采用思维链方式进行推理。这些高级推理能力为提升模型的安全性和鲁棒性开辟了新路径。特别是,当面对潜在不安全提示时,模型可在上下文中推理自身安全策略,实现深思熟虑的对齐。这使得其在生成非法建议、选择刻板回应及抵御已知越狱攻击等风险基准测试中达到当前最佳性能。在回答前引入思维链虽能带来显著收益,但也可能加剧因智能增强带来的潜在风险。结果强调需构建稳健的对齐方法,进行充分压力测试,并保持严谨的风险管理机制。本报告概述了针对OpenAI o1和OpenAI o1-mini模型开展的安全工作,包括安全评估、外部红队测试及准备框架评估。

原文摘要 · Abstract (English)

The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models. In particular, our models can reason about our safety policies in context when responding to potentially unsafe prompts, through deliberative alignment. This leads to state-of-the-art performance on certain benchmarks for risks such as generating illicit advice, choosing stereotyped responses, and succumbing to known jailbreaks. Training models to incorporate a chain of thought before answering has the potential to unlock substantial benefits, while also increasing potential risks that stem from heightened intelligence. Our results underscore the need for building robust alignment methods, extensively stress-testing their efficacy, and maintaining meticulous risk management protocols. This report outlines the safety work carried out for the OpenAI o1 and OpenAI o1-mini models, including safety evaluations, external red teaming, and Preparedness Framework evaluations.

推理增强安全对齐强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。