通过引导模型慢思考,提升安全防御能力。
Don't Command, Cultivate: An Exploratory Study of System-2 Alignment
- 用慢推理模式模拟人类系统2思维,增强安全判断。
- 模型对数学编码类越狱攻击仍存漏洞,需进一步优化。
- 简单提示工程可有效提升开源模型的安全对齐效果。
o1系列模型被认定为OpenAI中最稳健的模型,其核心特征是从快速直觉思维转向更慢、更审慎的推理。这一现象促使我们探究系统2思维模式对模型安全性的影响。初步研究中,我们对o1模型进行了安全评估,涵盖使用对抗性自然语言提示和数学编码提示的复杂越狱攻击场景。结果表明,o1模型表现出相对更好的安全性能,但仍易受数学编码类越狱攻击影响。通过对响应模式的详细案例分析,我们识别出特定行为模式。同时,我们探索了通过提示工程与监督微调在开源模型中实现系统2安全对齐的可行性。实验显示,一些简单方法能有效引导模型仔细审查用户请求,从而提升安全性。此外,我们提出了过程监督的实现方案,具体细节与实验结果将在后续版本中发布。
原文摘要 · Abstract (English)
The o1 system card identifies the o1 models as the most robust within OpenAI, with their defining characteristic being the progression from rapid, intuitive thinking to slower, more deliberate reasoning. This observation motivated us to investigate the influence of System-2 thinking patterns on model safety. In our preliminary research, we conducted safety evaluations of the o1 model, including complex jailbreak attack scenarios using adversarial natural language prompts and mathematical encoding prompts. Our findings indicate that the o1 model demonstrates relatively improved safety performance; however, it still exhibits vulnerabilities, particularly against jailbreak attacks employing mathematical encoding. Through detailed case analysis, we identified specific patterns in the o1 model's responses. We also explored the alignment of System-2 safety in open-source models using prompt engineering and supervised fine-tuning techniques. Experimental results show that some simple methods to encourage the model to carefully scrutinize user requests are beneficial for model safety. Additionally, we proposed a implementation plan for process supervision to enhance safety alignment. The implementation details and experimental results will be provided in future versions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。