通过临时越狱缓冲有害更新,实现大模型微调的安全与性能双赢。
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

- 临时越狱机制在微调中抑制有害梯度,保留有效任务梯度。
- 新框架在不依赖额外安全数据下,显著提升模型安全性与任务表现。
- 适合需要安全微调但资源受限的个性化AI应用开发者。
微调即服务(FaaS)使大语言模型个性化成为可能,但易受有害微调攻击导致安全对齐弱化。现有研究发现,在微调中激活有害行为模块可阻止模型学习不良行为,但机制尚不明确。本文重新审视临时越狱作为防御手段,通过梯度层面分析表明,其能饱和有害梯度同时保留良性任务相关梯度。基于此,提出缓冲与强化微调框架:BufferLoRA 在用户微调阶段引入可移除的越狱适配器以减少有害更新;适应后,通过QR分解融合训练恢复拒绝行为的ReinforceLoRA与UserLoRA,强化安全并保持用户任务性能。大量实验显示,该框架在无需额外安全数据、计算开销极小的前提下,实现更优的安全性与实用性。
原文摘要 · Abstract (English)
Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning attacks. Recent work has shown that activating harmful-behavior modules during fine-tuning can prevent models from learning undesired behaviors, but its mechanism remains unclear. In this paper, we revisit temporary jailbreaking as a defense against harmful fine-tuning and provide a gradient-level analysis showing that it saturates safety-degrading gradients while preserving benign task-relevant gradients. Based on this insight, we propose a Buffer-and-Reinforce fine-tuning framework that buffers harmful updates during user fine-tuning and reinforces safety after adaptation. Specifically, BufferLoRA induces temporary jailbreaking as a removable adapter to reduce harmful updates during user fine-tuning. After adaptation, ReinforceLoRA, trained to recover refusal behavior under the temporarily jailbroken state, is integrated with UserLoRA via QR decomposition-based merging to reinforce safety while preserving user-task performance. Extensive experiments show that our framework achieves superior safety and utility with no additional safety data during user fine-tuning and minimal computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。