arXiv:2509.16749cs.CRcs.AI2025-09中稿 · the Conference on …被引 11

评估大模型生成的网络安全规则效果,提升安全团队信任度。

Evaluating LLM Generated Detection Rules in Cybersecurity

  • 构建开源评估框架,对比大模型与人工编写的规则
  • 采用留出集方法,量化大模型规则的检测有效性
  • 提供专家视角的多维度指标,适合安全研发人员参考

大模型在安全领域的应用日益广泛,但其有效性缺乏有效评估手段,限制了安全从业者对其的信任与使用。本文提出一个开源评估框架和基准指标,用于评测大模型生成的网络安全规则。该基准采用留出集方法,将大模型生成的规则与人工编写的规则集进行对比,提供三个受专家评估方式启发的关键指标,实现对基于大模型的安全规则生成器的现实、多维度有效性评估。通过分析Sublime Security检测团队及自动化检测工程师(ADE)生成的规则,全面展示了ADE的能力表现。

原文摘要 · Abstract (English)

LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an open-source evaluation framework and benchmark metrics for evaluating LLM-generated cybersecurity rules. The benchmark employs a holdout set-based methodology to measure the effectiveness of LLM-generated security rules in comparison to a human-generated corpus of rules. It provides three key metrics inspired by the way experts evaluate security rules, offering a realistic, multifaceted evaluation of the effectiveness of an LLM-based security rule generator. This methodology is illustrated using rules from Sublime Security's detection team and those written by Sublime Security's Automated Detection Engineer (ADE), with a thorough analysis of ADE's skills presented in the results section.

大模型安全规则生成评估框架网络安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。