小模型易被越狱攻击,安全防护需从设计入手。
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
- 首次系统评估59个小型模型对12种越狱攻击的脆弱性。
- 超六成模型越狱成功率超40%,三成超50%。
- 防御效果依赖攻击类型,现有方法难以通用。
小型语言模型(SLMs)因计算开销低、隐私保护好且在特定领域表现接近大型模型,正成为边缘设备部署的新趋势。然而,其安全性尚未受到足够关注,尤其在越狱攻击方面。本文首次系统性地评估了59个来自15个主流家族的小型模型对12种前沿越狱攻击的脆弱性。结果显示,61.0%的模型在越狱攻击下平均攻击成功率(ASR)超过40%,37.3%的模型在直接有害查询下的ASR超过50%。相关性分析表明,模型脆弱性与训练细节(如训练数据集和方法)密切相关,而非模型规模。进一步测试五种防御策略发现,提示层防御在不同模型和攻击间表现不一致;模型层防御虽能提升对同类攻击的鲁棒性,但对多轮攻击(如Crescendo)泛化能力差,凸显了在小型模型开发中引入安全设计的紧迫性。
原文摘要 · Abstract (English)
Small language models (SLMs) have emerged as promising alternatives to large language models (LLMs) due to their low computational demands, enhanced privacy guarantees, and comparable performance in specific domains. Deploying SLMs on edge devices, such as smartphones and smart vehicles, has become a growing trend. However, the security implications of SLMs have not received as much attention as those of LLMs, particularly concerning the significant jailbreak threats they face. In this paper, we conduct the first systematic empirical study of SLMs' vulnerabilities to jailbreak attacks. Through systematic evaluation on 59 SLMs from 15 mainstream SLM families against 12 state-of-the-art jailbreak methods, we demonstrate that 61.0% of evaluated SLMs show an average ASR of more than 40% under jailbreak attacks and 37.3% of them have an ASR of more than 50% on direct harmful queries. Through correlation analysis, we identify that SLM vulnerabilities are closely related to training details (e.g., training dataset and method) rather than model size scaling. We further evaluate five defenses for jailbreak attacks, revealing that prompt-level defenses remain inconsistent across SLMs and attack methods, while model-level defense improves robustness against similar attacks yet generalizes poorly to multi-turn attacks such as Crescendo, highlighting the urgent need for security-by-design approaches in SLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。