arXiv:2502.19883cs.CRcs.AI2025-02ACL被引 2

小模型安全被忽视,实测发现其极易受越狱攻击

Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models

  • 对13个顶尖小模型进行越狱攻击实验
  • 多数小模型易被攻击,部分直接响应有害提示
  • 揭示压缩、量化等技术会降低模型安全性

小语言模型(SLMs)因高效低耗,日益广泛部署于边缘设备。尽管其性能通过创新训练和模型压缩不断进步,但相比大语言模型,其安全风险却长期未受重视。本文对13个最先进的SLMs在多种越狱攻击下进行全面实证研究,结果表明:大多数SLMs极易受到现有越狱攻击,部分甚至可被直接有害提示触发。为应对安全挑战,我们评估了若干典型防御方法并验证其有效性。进一步分析显示,架构压缩、量化、知识蒸馏等常见优化技术会引发潜在安全退化。本研究旨在揭示SLMs的安全隐患,为未来构建更鲁棒、更安全的小模型提供重要参考。

原文摘要 · Abstract (English)

Small language models (SLMs) have become increasingly prominent in the deployment on edge devices due to their high efficiency and low computational cost. While researchers continue to advance the capabilities of SLMs through innovative training strategies and model compression techniques, the security risks of SLMs have received considerably less attention compared to large language models (LLMs).To fill this gap, we provide a comprehensive empirical study to evaluate the security performance of 13 state-of-the-art SLMs under various jailbreak attacks. Our experiments demonstrate that most SLMs are quite susceptible to existing jailbreak attacks, while some of them are even vulnerable to direct harmful prompts.To address the safety concerns, we evaluate several representative defense methods and demonstrate their effectiveness in enhancing the security of SLMs. We further analyze the potential security degradation caused by different SLM techniques including architecture compression, quantization, knowledge distillation, and so on. We expect that our research can highlight the security challenges of SLMs and provide valuable insights to future work in developing more robust and secure SLMs.

小模型越狱攻击安全边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。