评估大模型推理安全风险,发现开源模型易被攻击且思考过程更危险
The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1
- 对比多个推理模型的安全表现与抗攻击能力
- 更强推理能力带来更大潜在危害,尤其在应对不当问题时
- 模型思考过程比最终回答更需关注安全,适合安全研究者参考
大型推理模型(LRMs)如 OpenAI-o3 和 DeepSeek-R1 在复杂推理任务上显著优于传统大语言模型(LLMs)。然而,其强大能力与开源模型(如 DeepSeek-R1)的广泛可及性,引发严重安全担忧。本文通过主流安全基准评估模型合规性,并测试其对越狱攻击和提示注入等对抗性攻击的脆弱性。多维度分析揭示四大发现:(1) 开源推理模型与 o3-mini 模型在安全基准和抗攻击能力上存在显著差距,表明需加强开源模型的安全投入;(2) 模型推理能力越强,回答不当问题时造成的潜在危害越大;(3) 虽然模型在推理中展现出安全意识,但常被对抗攻击突破;(4) 模型的推理过程本身比最终输出更易引发安全风险。研究揭示了推理模型的安全隐患,强调需进一步提升 R1 类模型的安全性以缩小差距。
原文摘要 · Abstract (English)
The rapid development of large reasoning models (LRMs), such as OpenAI-o3 and DeepSeek-R1, has led to significant improvements in complex reasoning over non-reasoning large language models~(LLMs). However, their enhanced capabilities, combined with the open-source access of models like DeepSeek-R1, raise serious safety concerns, particularly regarding their potential for misuse. In this work, we present a comprehensive safety assessment of these reasoning models, leveraging established safety benchmarks to evaluate their compliance with safety regulations. Furthermore, we investigate their susceptibility to adversarial attacks, such as jailbreaking and prompt injection, to assess their robustness in real-world applications. Through our multi-faceted analysis, we uncover four key findings: (1) There is a significant safety gap between the open-source reasoning models and the o3-mini model, on both safety benchmark and attack, suggesting more safety effort on open LRMs is needed. (2) The stronger the model's reasoning ability, the greater the potential harm it may cause when answering unsafe questions. (3) Safety thinking emerges in the reasoning process of LRMs, but fails frequently against adversarial attacks. (4) The thinking process in R1 models poses greater safety concerns than their final answers. Our study provides insights into the security implications of reasoning models and highlights the need for further advancements in R1 models' safety to close the gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。