arXiv:2506.18543cs.CRcs.AI2025-06被引 3

首次全面评估DeepSeek模型的越狱防御能力,发现其安全性不如GPT-4。

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models

  • 对比GPT-3.5、GPT-4与DeepSeek在510种有害行为上的越狱攻击效果
  • DeepSeek对优化型攻击有部分防御力,但对提示工程类攻击更脆弱
  • 适合关注开源大模型安全性的研究人员和部署者参考

大型语言模型(LLMs)的快速普及引发了对其越狱攻击风险的关注,此类攻击通过构造对抗性输入诱导模型输出不安全内容。尽管闭源模型如GPT-4已得到广泛评估,但新兴开源模型如DeepSeek的安全性仍缺乏充分研究。本文首次对DeepSeek模型族进行系统性越狱分析,通过HarmBench基准与GPT-3.5、GPT-4对比,考察七种代表性攻击方法在510种有害行为上的表现。结果表明,DeepSeek对优化驱动型攻击(如TAP-T)具有部分鲁棒性,但对提示工程及人工设计的对抗输入更易受攻击。相比之下,GPT-4 Turbo在多种行为上展现出更强且更一致的安全对齐,可能归因于更优的安全优化与基于人类反馈的强化学习。细粒度行为分析与案例研究显示,DeepSeek常无法一致执行安全约束,导致拒绝行为不一致。总体而言,研究揭示了模型效率与对齐泛化之间的内在权衡,强调需针对性地开展安全调优与鲁棒对齐策略,以保障开源大模型的安全部署。

原文摘要 · Abstract (English)

The rapid proliferation of Large Language Models (LLMs) has heightened concerns regarding their exposure to jailbreak attacks, which craft adversarial inputs designed to elicit unsafe content. Although proprietary models such as GPT-4 have been extensively evaluated, the robustness of emerging open-source systems like DeepSeek remains insufficiently examined, despite their growing use in LLM applications. In this paper, we conduct the first comprehensive jailbreak analysis of the DeepSeek model family, comparing it with GPT-3.5 and GPT-4 through the HarmBench benchmark. We investigate seven representative attack methods across 510 harmful behaviors, organized along both functional and semantic dimensions. Findings indicate that DeepSeek provides partial resilience against optimization-driven attacks such as TAP-T, but also results in greater susceptibility to prompt-based and manually engineered adversarial inputs. In contrast, GPT-4 Turbo demonstrates more robust and consistent safety alignment across a wide range of behaviors, likely due to stronger safety optimization and reinforcement learning from human feedback. In addition, fine-grained behavioral analysis and case studies reveal that DeepSeek often fails to consistently apply safety constraints to adversarial prompts, leading to uneven refusal behaviors. Overall, our results highlight an inherent trade-off between model efficiency and alignment generalization, underscoring the importance of targeted safety tuning and robust alignment strategies to ensure secure deployment of open-source LLMs.

大模型安全越狱攻击开源模型对抗测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。