arXiv:2506.00166cs.LGcs.AI2025-06被引 5

分离式安全模块让模型更安全、更高效,还能灵活调节安全等级。

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

  • 将安全功能与主模型解耦,用轻量适配器实现动态安全控制。
  • 在多个任务上性能提升最高达53%的AUC,且推理开销极低。
  • 支持运行时调整安全强度,适合需要灵活安全策略的部署场景。

现有AI安全范式如防护模型和对齐训练常牺牲推理效率或开发灵活性。本文提出解耦安全适配器(DSA),通过将安全计算从任务优化的基模型中分离,实现高效且灵活的安全机制。DSA采用轻量适配器,利用基模型内部表征,以最小推理成本支持多样化安全功能。实验表明,基于DSA的安全防护在仇恨言论分类、不安全输入/输出检测及幻觉识别任务中,相较同等规模独立模型,AUC最高提升53%。此外,基于DSA的安全对齐可实现运行时动态调节对齐强度,精细平衡指令遵循与安全性。结合安全防护与对齐机制后,对StrongREJECT的防护能力提升93%,同时保持MTBench上98%的性能,相比标准对齐微调减少8个百分点的对齐代价。整体上,DSA为更模块化、高效、可扩展的AI安全与对齐提供了新路径。

原文摘要 · Abstract (English)

Existing paradigms for ensuring AI safety, such as guardrail models and alignment training, often compromise either inference efficiency or development flexibility. We introduce Disentangled Safety Adapters (DSA), a novel framework addressing these challenges by decoupling safety-specific computations from a task-optimized base model. DSA utilizes lightweight adapters that leverage the base model's internal representations, enabling diverse and flexible safety functionalities with minimal impact on inference cost. Empirically, DSA-based safety guardrails substantially outperform comparably sized standalone models across hate speech classification, detecting unsafe model inputs and responses, and hallucination detection with relative improvements of up to 53% in AUC. Furthermore, DSA-based safety alignment allows dynamic, inference-time adjustment of alignment strength and a fine-grained trade-off between instruction following performance and model safety. Importantly, combining the DSA safety guardrail with DSA safety alignment facilitates context-dependent alignment strength, boosting safety on StrongREJECT by 93% while maintaining 98% performance on MTBench - a total reduction in alignment tax of 8 percentage points compared to standard safety alignment fine-tuning. Overall, DSA presents a promising path towards more modular, efficient, and adaptable AI safety and alignment.

AI安全适配器灵活对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。