量化大模型可能不安全,该研究提出修复方案恢复其安全性。
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
- 针对量化后大模型的安全风险,设计了感知量化的安全修复框架。
- 在多种量化方法下,修复后模型安全性接近原始未量化版本。
- 适合关注模型部署安全性的研究人员与工程团队。
量化大语言模型(LLMs)因能在资源受限环境下部署而日益受到重视。然而,近期一些无需校准数据集的量化方法研究指出,量化可能损害LLM的安全能力,凸显出系统性安全评估与有效缓解策略的迫切需求。本文对主流量化技术及多种校准数据集进行了全面的安全评估,采用广泛认可的安全基准测试。为应对发现的安全漏洞,我们提出一种量化感知的安全修复框架Q-resafe,可高效恢复量化后LLM的安全能力,同时最小化对模型性能的负面影响。大量实验表明,即使在挑战性评估场景下,Q-resafe仍能将量化模型的安全性恢复至接近原始未量化水平。项目主页见:https://github.com/Thecommonirin/Qresafe。
原文摘要 · Abstract (English)
Quantized large language models (LLMs) have gained increasing attention and significance for enabling deployment in resource-constrained environments. However, emerging studies on a few calibration dataset-free quantization methods suggest that quantization may compromise the safety capabilities of LLMs, underscoring the urgent need for systematic safety evaluations and effective mitigation strategies. In this paper, we present comprehensive safety evaluations across various mainstream quantization techniques and diverse calibration datasets, utilizing widely accepted safety benchmarks. To address the identified safety vulnerabilities, we propose a quantization-aware safety patching framework, Q-resafe, to efficiently restore the safety capabilities of quantized LLMs while minimizing any adverse impact on utility. Extensive experimental results demonstrate that Q-resafe successfully re-aligns the safety of quantized LLMs with their pre-quantization counterparts, even under challenging evaluation scenarios. Project page is available at: https://github.com/Thecommonirin/Qresafe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。