构建首个网络安全领域专用的攻击提示数据集,用于精准评测大模型越狱能力。
CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
- 用生成式AI构建1.2万条网络安全定向攻击提示,分10类且为封闭式问题
- 在商业大模型上测试越狱成功率:ChatGPT达65%,Gemini达88%,Claude仅17%
- 适合安全研究人员、模型开发者评估和防御越狱攻击
大量研究探讨了如何使大型语言模型(LLMs)生成有害内容。通常,这些方法通过设计恶意提示的数据集来评估其绕过模型提供方安全策略的能力。然而,现有数据集普遍范围宽泛、开放性强,难以在特定领域(如网络安全)准确评估越狱效果。为此,我们提出并公开发布CySecBench,一个包含12662个专为网络安全领域设计的提示数据集,涵盖10种不同攻击类型,采用封闭式问题以实现更一致、精确的越狱评估。此外,我们详细描述了数据集生成与过滤方法,可推广至其他领域。为验证其价值,我们提出一种基于提示混淆的越狱方法,在商用黑盒模型上取得65%(ChatGPT)、88%(Gemini)的成功率;而Claude表现出更强韧性,成功率为17%。相比现有基准方法,本方法表现更优。在使用通用数据集AdvBench测试时,成功率达78.5%,高于当前最先进水平。
原文摘要 · Abstract (English)
Numerous studies have investigated methods for jailbreaking Large Language Models (LLMs) to generate harmful content. Typically, these methods are evaluated using datasets of malicious prompts designed to bypass security policies established by LLM providers. However, the generally broad scope and open-ended nature of existing datasets can complicate the assessment of jailbreaking effectiveness, particularly in specific domains, notably cybersecurity. To address this issue, we present and publicly release CySecBench, a comprehensive dataset containing 12662 prompts specifically designed to evaluate jailbreaking techniques in the cybersecurity domain. The dataset is organized into 10 distinct attack-type categories, featuring close-ended prompts to enable a more consistent and accurate assessment of jailbreaking attempts. Furthermore, we detail our methodology for dataset generation and filtration, which can be adapted to create similar datasets in other domains. To demonstrate the utility of CySecBench, we propose and evaluate a jailbreaking approach based on prompt obfuscation. Our experimental results show that this method successfully elicits harmful content from commercial black-box LLMs, achieving Success Rates (SRs) of 65% with ChatGPT and 88% with Gemini; in contrast, Claude demonstrated greater resilience with a jailbreaking SR of 17%. Compared to existing benchmark approaches, our method shows superior performance, highlighting the value of domain-specific evaluation datasets for assessing LLM security measures. Moreover, when evaluated using prompts from a widely used dataset (i.e., AdvBench), it achieved an SR of 78.5%, higher than the state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。