arXiv:2502.18935cs.CLcs.AI2025-02KDD被引 9

首个针对中文大模型安全漏洞的综合评估基准,能有效发现潜在风险。

JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models

  • 构建基于中文语境的分层安全分类体系,精准定位漏洞。
  • 使用自动越狱提示生成框架,提升测试效率与覆盖面。
  • 在13个主流模型上验证,对ChatGPT攻击成功率最高,适合安全研究者使用。

大语言模型(LLMs)在各类应用中展现出卓越能力,亟需全面的安全评估。尤其随着中文语言能力增强及汉语表达的独特性,针对中文场景的安全评估基准应运而生。然而,现有基准普遍难以有效暴露模型安全漏洞。为此,我们提出JailBench——首个面向中文大模型深层漏洞的综合性评估基准,包含专为中文设计的精细化分层安全分类体系。为提升生成效率,采用创新的自动越狱提示工程师(AJPE)框架,融合越狱技术以增强评估效果,并利用大模型通过上下文学习自动扩展数据集。JailBench在13个主流大模型上进行广泛评测,对ChatGPT的攻击成功率高于现有中文基准,充分验证其识别潜在漏洞的能力,揭示中文场景下大模型安全性仍有巨大提升空间。基准已开源:https://github.com/STAIR-BUPT/JailBench。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities across various applications, highlighting the urgent need for comprehensive safety evaluations. In particular, the enhanced Chinese language proficiency of LLMs, combined with the unique characteristics and complexity of Chinese expressions, has driven the emergence of Chinese-specific benchmarks for safety assessment. However, these benchmarks generally fall short in effectively exposing LLM safety vulnerabilities. To address the gap, we introduce JailBench, the first comprehensive Chinese benchmark for evaluating deep-seated vulnerabilities in LLMs, featuring a refined hierarchical safety taxonomy tailored to the Chinese context. To improve generation efficiency, we employ a novel Automatic Jailbreak Prompt Engineer (AJPE) framework for JailBench construction, which incorporates jailbreak techniques to enhance assessing effectiveness and leverages LLMs to automatically scale up the dataset through context-learning. The proposed JailBench is extensively evaluated over 13 mainstream LLMs and achieves the highest attack success rate against ChatGPT compared to existing Chinese benchmarks, underscoring its efficacy in identifying latent vulnerabilities in LLMs, as well as illustrating the substantial room for improvement in the security and trustworthiness of LLMs within the Chinese context. Our benchmark is publicly available at https://github.com/STAIR-BUPT/JailBench.

大模型安全中文评估越狱检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。