安全适配器SafeGene可反复使用,修复大模型微调后的安全漏洞。
SafeGene: Reusable Adapters for Transferable Safety Alignment

- 将安全能力作为独立适配模块,脱离任务更新
- 减少有害回复率,同时保持下游任务性能
- 适合需要持续安全维护的定制化大模型应用
开放权重的大语言模型正被不断微调为个性化助手,但下游微调可能削弱其安全对齐能力,使模型更易受恶意提示影响,即使训练数据无恶意。这导致安全问题反复出现,因目标模型频繁更新新任务数据或用户交互。我们提出SafeGene,一种可在架构兼容模型家族中跨任务复用的安全适配模块。不同于将安全修复视为模型特异的修补步骤,SafeGene将安全能力视为与任务无关、可复用的适配表示,通过对齐-退化模型差异提取,结合数据感知层选择优化为任务可迁移的安全向量,并以少量样本进行逐层系数重校准,嵌入每个下游任务适配模型。多模型族、多任务及多个安全评估者实验表明,SafeGene增强模型在保持下游性能的同时显著降低有害回复率,在安全-效用权衡上优于现有主流安全适配方法。
原文摘要 · Abstract (English)
Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vulnerable to malicious prompts, even when the training data is not intentionally harmful. This creates a recurring safety recovery problem as target models are repeatedly updated with new task data or user interactions. We propose SafeGene, a reusable safety-adapter module designed for cross-task reuse within each architecture-compatible model family. Rather than treating safety recovery as a model-specific repair step, SafeGene treats safety capability as an independent, reusable adapter representation decoupled from task-specific updates. This representation is obtained from aligned--degraded model discrepancies, refined into task-transferable safety vectors through data-aware layer selection, and expressed in each downstream task-adapted model via few-shot layer-wise coefficient recalibration. Experiments across multiple model families, downstream tasks, and safety judges show that SafeGene-enhanced models reduce harmful response rates while maintaining downstream performance, outperforming representative safe adaptation methods in safety--utility trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。