首次系统评估大模型越狱防护机制,为安全部署提供标准框架。
SoK: Evaluating Jailbreak Guardrails for Large Language Models
- 构建六维分类体系,统一梳理越狱防护手段
- 提出安效用评估框架,量化防护效果与性能代价
- 揭示现有方案在多攻击类型下的通用性短板
大语言模型(LLMs)虽取得显著进展,但其部署暴露了越狱攻击等关键漏洞,此类攻击可绕过安全对齐。外部防护机制——即监控和控制LLM交互的守卫机制(guardrails)——成为有前景的解决方案。然而,当前的守卫机制研究分散,缺乏统一分类与全面评估框架。本文作为知识体系化(SoK)论文,首次对大模型越狱防护机制进行整体分析,提出一个六维多维度分类体系,并引入安全-效率-效用评估框架以衡量实际有效性。通过广泛分析与实验,我们识别出现有方法的优势与局限,提供优化防御机制的洞见,并探索其在多种攻击类型中的普适性。本工作为未来研究与开发奠定结构化基础,旨在推动鲁棒守卫机制的规范演进与部署。代码已开源:https://github.com/xunguangwang/SoK4JailbreakGuardrails。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable progress, but their deployment has exposed critical vulnerabilities, particularly to jailbreak attacks that circumvent safety alignments. Guardrails--external defense mechanisms that monitor and control LLM interactions--have emerged as a promising solution. However, the current landscape of LLM guardrails is fragmented, lacking a unified taxonomy and comprehensive evaluation framework. In this Systematization of Knowledge (SoK) paper, we present the first holistic analysis of jailbreak guardrails for LLMs. We propose a novel, multi-dimensional taxonomy that categorizes guardrails along six key dimensions, and introduce a Security-Efficiency-Utility evaluation framework to assess their practical effectiveness. Through extensive analysis and experiments, we identify the strengths and limitations of existing guardrail approaches, provide insights into optimizing their defense mechanisms, and explore their universality across attack types. Our work offers a structured foundation for future research and development, aiming to guide the principled advancement and deployment of robust LLM guardrails. The code is available at https://github.com/xunguangwang/SoK4JailbreakGuardrails.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。