系统化梳理大模型提示安全,统一评估攻击与防御方法。
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- 构建攻击、防御与漏洞的关联分类体系,分离技术机制与能力假设。
- 实证发现:攻击成本、评估者选择、防御反噬等显著影响安全结论。
- 开源工具链支持可复现评估,适合研究者和安全测试人员使用。
大型语言模型正广泛用于信息、代码及现实服务接口,提示层安全缺陷已成为实际风险。尽管越狱攻击、防御手段、数据集与自动评判工具快速发展,评估仍分散于威胁模型、访问假设、成本预算、数据集与成功标准之间,导致报告的攻击成功率与防御提升难以比较。本文系统化整理大模型提示安全的范畴、数据、工具与度量方式。提出越狱攻击、防御与模型漏洞的关联分类体系,将技术机制与攻防能力分离。形式化定义威胁、访问与成本假设为显式评估元数据。为支持可复现评估,发布 JailbreakDB、PromptSecurity-Eval 及 PromptSecurity 平台,将每项实验表示为模型、攻击、防御、数据集与评判器的组合元组。通过跨模型、攻击、防御与评判器的匹配评估,表明访问模式、原生有害查询行为、攻击成本、防御反噬、分类子类与评判器选择均显著影响安全结论。上述成果共同支撑可复现、成本敏感且基于分类体系的大模型提示安全评估。排行榜:https://datasec-lab.github.io/PromptSecurityLeaderboard/。数据集:https://huggingface.co/datasets/youbin2014/JailbreakDB。GitHub:https://github.com/datasec-lab/PromptSecurity。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used as interfaces to information, code, and real-world services, making prompt-level security failures a practical concern. Although jailbreak attacks, defenses, datasets, and automated judgers have advanced rapidly, evaluation remains fragmented across threat models, access assumptions, cost budgets, datasets, and success criteria. This makes reported attack success rates and defense gains hard to compare. This SoK systematizes LLM prompt security across concepts, data, tooling, and measurement. We propose linked taxonomies for jailbreak attacks, defenses, and model vulnerabilities, while separating technical mechanisms from attacker and defender capabilities. We also formalize threat, access, and cost assumptions as explicit evaluation metadata. To support reproducible evaluation, we release JailbreakDB, PromptSecurity-Eval, and PromptSecurity, a modular platform that represents each experiment as a tuple of model, attack, defense, dataset, and judger. Using matched evaluations across models, attacks, defenses, and judgers, we show that access regime, native harmful-query behavior, attack cost, defense backfire, taxonomy subcategory, and judger choice all materially affect security conclusions. Together, these artifacts support reproducible, cost-aware, and taxonomy-grounded evaluation of LLM prompt security. Leaderboard: https://datasec-lab.github.io/PromptSecurityLeaderboard/. Dataset: https://huggingface.co/datasets/youbin2014/JailbreakDB. GitHub: https://github.com/datasec-lab/PromptSecurity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。