arXiv:2412.05346cs.CRcs.LG2024-12

仅用简单微调即可移除GPT-4o安全防护,且不损失性能。

BadGPT-4o: stripping safety finetuning from GPT models

  • 通过微调污染技术剥离模型安全机制
  • 在HarmBench和StrongREJECT上媲美顶尖越狱攻击
  • 无额外令牌开销,适合快速部署

我们展示了一种基于Qi等人2023年提出的简单微调污染技术的BadGPT-4o攻击,可有效移除GPT-4o的安全防护机制,且不降低模型性能。该攻击在HarmBench和StrongREJECT测试中表现与最佳白盒越狱攻击相当。在tinyMMLU和开放生成任务中,其无需额外令牌开销,也未引发性能下降。尽管此攻击已公开一年,但其执行依然简便,存在显著安全风险。

原文摘要 · Abstract (English)

We show a version of Qi et al. 2023's simple fine-tuning poisoning technique strips GPT-4o's safety guardrails without degrading the model. The BadGPT attack matches best white-box jailbreaks on HarmBench and StrongREJECT. It suffers no token overhead or performance hits common to jailbreaks, as evaluated on tinyMMLU and open-ended generations. Despite having been known for a year, this attack remains easy to execute.

越狱攻击模型安全微调污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。