激活近似会削弱对齐大模型的安全性,提出新方法防御此风险。
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
- 分析七种激活近似技术的安全隐患
- 发现十款对齐模型安全性能普遍下降
- 提出QuadA方法提升近似后模型安全性
大型语言模型(LLMs)在多个领域展现出卓越能力。随着其能力提升和部署场景扩展,模型规模庞大及激活设计复杂带来的部署挑战日益突出,尤其在资源受限场景中,缓解推理瓶颈至关重要。近年来,激活近似成为提升推理效率的重要手段,常被视为私有推理等场景的必要选项。尽管其在保持性能的同时实现显著加速,但激活近似对安全性的潜在影响仍不明确。本文首次系统评估了激活近似在安全方面的风险,覆盖三种主流类别(激活多项式化、稀疏化、量化)下的七种前沿技术,发现在十款安全对齐的LLM上均出现持续的安全退化。为应对多种近似方法共存的防御难题,我们深入分析其共享误差模式,得出三个关键发现,并提出针对该问题的新型安全增强方法QuadA。大量实验与消融研究验证了QuadA在激活近似后有效提升模型安全性的能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have showcased remarkable capabilities across various domains. Accompanying the evolving capabilities and expanding deployment scenarios of LLMs, their deployment challenges escalate due to their sheer scale and the advanced yet complex activation designs prevalent in notable model series, such as Llama, Gemma, Mistral. These challenges have become particularly pronounced in resource-constrained deployment scenarios, where mitigating inference bottlenecks is imperative. Among various recent efforts, activation approximation has emerged as a promising avenue for pursuing inference efficiency, sometimes considered indispensable in applications such as private inference. Despite achieving substantial speedups with minimal impact on utility, even appearing sound and practical for real-world deployment, the safety implications of activation approximations remain unclear. In this work, we fill this critical gap in LLM safety by conducting the first systematic safety evaluation of activation approximations. Our safety vetting spans seven state-of-the-art techniques across three popular categories (activation polynomialization, activation sparsification, and activation quantization), revealing consistent safety degradation across ten safety-aligned LLMs. To overcome the hurdle of devising a unified defense accounting for diverse activation approximation methods, we perform an in-depth analysis of their shared error patterns and uncover three key findings. We propose QuadA, a novel safety enhancement method tailored to mitigate the safety compromises introduced by activation approximations. Extensive experiments and ablation studies corroborate QuadA's effectiveness in enhancing the safety capabilities of LLMs after activation approximations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。