仅用一个无标签提示,就能让大模型失去安全约束。
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
- 用群体相对策略优化法,直接移除模型安全限制。
- 单个无标签提示即可可靠使15个大模型失衡,且保持功能可用。
- 适用语言与图像生成模型,适合研究模型安全弱点者。
安全对齐的鲁棒性取决于其最弱的失效模式。尽管已有大量安全后训练工作,但研究表明模型可通过部署后微调快速失衡。现有方法常需大量数据整理且损害模型实用性。本文提出GRP-Obliteration(GRP-Oblit),利用组相对策略优化(GRPO)直接解除目标模型的安全约束。实验表明,仅需一个无标签提示即可可靠使安全对齐模型失衡,同时基本保持其功能性能。相比现有最先进方法,GRP-Oblit在平均意义上实现更强的失衡效果。该方法还拓展至基于扩散的图像生成系统。我们在六项通用性基准与五项安全性基准上,评估了十五个7-200亿参数规模的模型,涵盖GPT-OSS、DeepSeek蒸馏版、Gemma、Llama、Ministral及Qwen等指令型与推理型模型,包括密集与MoE架构。
原文摘要 · Abstract (English)
Safety alignment is only as robust as its weakest failure mode. Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-deployment fine-tuning. However, these methods often require extensive data curation and degrade model utility. In this work, we extend the practical limits of unalignment by introducing GRP-Obliteration (GRP-Oblit), a method that uses Group Relative Policy Optimization (GRPO) to directly remove safety constraints from target models. We show that a single unlabeled prompt is sufficient to reliably unalign safety-aligned models while largely preserving their utility, and that GRP-Oblit achieves stronger unalignment on average than existing state-of-the-art techniques. Moreover, GRP-Oblit generalizes beyond language models and can also unalign diffusion-based image generation systems. We evaluate GRP-Oblit on six utility benchmarks and five safety benchmarks across fifteen 7-20B parameter models, spanning instruct and reasoning models, as well as dense and MoE architectures. The evaluated model families include GPT-OSS, distilled DeepSeek, Gemma, Llama, Ministral, and Qwen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。