arXiv:2410.16527cs.CRcs.LG2024-10被引 15

对比主流开源LLM安全扫描工具,揭示检测可靠性短板

Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis

  • 采用红队测试思路评估四款开源扫描器的漏洞发现能力
  • 实测显示多数工具对成功攻击检测率不足60%,存在明显误判
  • 适合安全团队选型参考,尤其关注定制化与行业适配性

本报告对对话式大语言模型(LLMs)的开源漏洞扫描工具进行对比分析。随着LLMs广泛应用于各类场景,其面临信息泄露和越狱攻击等安全风险。研究评估了Garak、Giskard、PyRIT和CyberSecEval四款主流扫描器,它们均借鉴红队实践以暴露潜在漏洞。文章详述各工具特性与使用方法,总结设计共性,并开展量化评估。结果表明,这些工具在识别成功攻击方面存在显著可靠性问题,暴露出当前技术核心缺口。此外,研究贡献了一个初步标注的数据集,为后续研究提供基础。基于分析,提出组织在选择扫描器时应综合考虑可定制性、测试用例全面性及行业应用场景。

原文摘要 · Abstract (English)

This report presents a comparative analysis of open-source vulnerability scanners for conversational large language models (LLMs). As LLMs become integral to various applications, they also present potential attack surfaces, exposed to security risks such as information leakage and jailbreak attacks. Our study evaluates prominent scanners - Garak, Giskard, PyRIT, and CyberSecEval - that adapt red-teaming practices to expose these vulnerabilities. We detail the distinctive features and practical use of these scanners, outline unifying principles of their design and perform quantitative evaluations to compare them. These evaluations uncover significant reliability issues in detecting successful attacks, highlighting a fundamental gap for future development. Additionally, we contribute a preliminary labelled dataset, which serves as an initial step to bridge this gap. Based on the above, we provide strategic recommendations to assist organizations choose the most suitable scanner for their red-teaming needs, accounting for customizability, test suite comprehensiveness, and industry-specific use cases.

LLM安全漏洞扫描红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。